What is measured, where each figure comes from, and the terms its source publishes it under. Glototeca's own downloads are not available yet; every original source is, and each one is linked here.
What we mean by digital resources
This means the technical building blocks for language tools: translators, speech recognition, spell-checkers, predictive keyboards. It does not mean cultural content such as books, music, video or websites written in the language.
The distinction matters because they are different things. A language can have a large literature and no corpus a model could be trained on; the reverse also happens. This site measures only the second, and says nothing about a language's cultural vitality.
The five kinds that are counted
Datasets
Collections of text or audio prepared for a program to read.
Models
Trained systems that declare they work with the language.
Speech corpora
Recordings with their transcriptions, the basis of speech recognition.
Treebanks
Sentences with their grammatical analysis annotated by hand.
Speech recognition support
Whether a recognition model lists the language among those it supports.
They allow copying, distributing, adapting and commercial use. They require credit as "Fuente: INEGI, [product]" and disclosure of any transformation, without presenting it as done or endorsed by INEGI.
The document, published in the Diario Oficial de la Federación on 14 January 2008, contains no rights notice, licence or reproduction terms. All 256 pages were checked.
The legal page prohibits reproduction in whole or in part without written permission.
The risk grades shown on this site were extracted from this all-rights-reserved publication. Permission to redistribute them as open data, with attribution, will be requested from INALI; it has not yet been asked for or confirmed.
Original source
Go to the sourcesite.inali.gob.mx/pdf/libro_lenguas_indigenas_nacionales_en_riesgo_de_desaparicion.pdf
Glototeca's processed version
Coming soon The download does not exist yet.
A note on licences
Two things with different terms sit side by side here.
Glototeca's own work is published under CC BY 4.0: the crosswalk between groups, ISO 639-3 codes and Glottocodes; the weekly measurements with their four states; the variant-level resolution; and the findings on direction errors in variant names. It will be downloadable from this page once that feature exists.
Data drawn from third parties keeps whatever licence each source publishes, shown above with the date it was checked. Glototeca does not claim the right to redistribute it in bulk until that is confirmed source by source.
One case is open. The risk grades come from a 2012 INALI book that reserves all rights. Their status is not settled either way: they are not published under CC BY 4.0 as if it were, and they depend on INALI's answer. The current status is on that source's card.