Method
What Glototeca measures, where the data comes from, how we decide that two names are the same language, and what we do not know yet.
What this measures and what it does not
Once a week, Glototeca records which public digital resources exist for each of the 68 language groups in INALI's catalog: datasets, models, speech corpora, treebanks and support in one speech-recognition tool. It sets them beside two pieces of context: how many people speak the language and how much at risk it is.
It measures presence in public sources, not quality or usefulness. A dataset existing for a language says nothing about whether it is good, whether the community knows about it or whether it serves them. And nothing showing up does not mean nothing exists: it means nothing shows up in the sources we check.
The four states of a value
Ayapaneco: 3 datasets on Hugging Face carry its tag. →
Ayapaneco: we checked Universal Dependencies, which could hold a treebank for any language, and there is none. This is a real zero. →
Ayapaneco: Common Voice publishes a closed list of languages and this one is not on it. That is not a zero: the source does not include it. →
Ku’ahl: it has no ISO 639-3 code, so there is nothing to look it up by in any source. The value is unknown, and is shown as unknown. →
No single score
The indicators are never combined into a rating. Adding datasets to treebanks, or weighting one against another, would produce a number that looks precise and means nothing. Each indicator is shown separately, with its source and its date.
Mexico is the pilot. The long-term aim is to apply the same method to every language in Glottolog's reference list (7,674 spoken languages in version 5.3); today the site covers Mexico only.
Where the data comes from
Eight sources, all public and free. Five are checked every week; three are fixed publications that were loaded once. The latest snapshot is from Oct 5, 2026.
| Source | What it provides | How it is obtained | Cadence | Version on record |
|---|---|---|---|---|
| Glottolog | Each language's identity and its endangerment status | A dated release, downloaded from Zenodo and pinned by its DOI | Weekly | glottolog 5.3 (doi:10.5281/zenodo.18840967) |
| Hugging Face | Datasets and models tagged with the language | Public API, no account | Weekly | Oct 5, 2026 |
| Common Voice | Whether a speech corpus exists, and for how many varieties | Metadata for each release, published on GitHub | Weekly | common-voice scripted 27.0 (2026-09-11), spontaneous 5.0 (2026-09-11) |
| Universal Dependencies | Released treebanks | Twice-yearly releases on GitHub | Weekly | universal-dependencies 2.18 (2026-05-15) |
| Omnilingual ASR | Whether the speech-recognition model declares support | The model's own language list, by code | Weekly | omnilingual-asr lang_ids.py @ a7fb36017a (2025-11-10) |
| INEGI, 2020 Census | Speakers aged 3 and over, by language | Official census table, downloaded once | Fixed | INEGI Censo de Población y Vivienda 2020, cpv2020_b_eum_05_etnicidad.xlsx sheet 03 |
| INALI, Catálogo de las Lenguas Indígenas Nacionales (2008) | The 364 variants: names, autonyms and localities | Official PDF, read by a program and verified by its checksum | Fixed | |
| INALI, Lenguas indígenas nacionales en riesgo de desaparición (2012) | The risk grade of each variant | Official PDF, read by a program and verified by its checksum | Fixed |
No source requires payment or a key. What a source publishes under its own terms stays under those terms.
How identity is resolved
Every source names languages its own way. INALI speaks of groups and variants; Hugging Face and Common Voice use ISO 639-3 codes; Glottolog uses its own identifiers. Before counting anything, someone has to decide which code belongs to which group. That table of correspondences was built once, by hand, and checked against Glottolog's own cross-references. Everything else reads it automatically.
Of the 68 groups, 65 are resolved: 43 correspond to a single Glottolog language, 21 bring several together and 1 is a macrolanguage. The remaining 3 (ku'ahl, otomí, zoque) are not resolved, and the site shows them that way. A doubtful identity is flagged; it is never guessed.
An error we found
The 2008 catalog and the 2012 risk book are two INALI publications about the same 364 variants. When we joined them, 354 names matched letter for letter. Among those that did not, 3 differ by a single word, and that word is a direction: where one publication says "noreste" (northeast), the other says "noroeste" (northwest).
This is not a spelling variant. They are two different directions. What shows that each pair is one and the same variant is a different kind of evidence: the number of localities each publication assigns to it.
| 2008 catalog | 2012 risk book | Localities (2008) | Localities (2012) |
|---|---|---|---|
| zapoteco de la Sierra sur, noroeste | zapoteco de la Sierra sur, noreste | 16 | 16 |
| zapoteco de la Sierra sur, noreste alto | zapoteco de la Sierra sur, noroeste alto | 13 | 13 |
| náhuatl del noreste central | náhuatl del noroeste central | 400 | 398 |
The error is most likely in the 2008 catalog, though we cannot rule out that it is in the 2012 book. The documents alone do not settle it.
So we did not correct it. The site keeps the name exactly as the catalog prints it, links each variant to its row in the risk book, and says on the variant's page how the link was made. The discrepancy stays flagged in the data, in plain view, so that INALI or any specialist can resolve it.
Limitations
What to know before quoting a figure from this site.
- Risk and speakers come from different censuses
- The risk grades were calculated from the 2000 census. The speaker counts we show are from the 2020 census. Twenty years separate them: a variant may be better or worse off today than its grade says.
- Most variants have no identifier of their own
- Only 136 of the 364 variants carry a Glottocode, and it is their group's, not their own. The other 228 belong to groups that correspond to several Glottolog languages, and there is no reliable way to assign them. That is why digital resources are measured per group, not per variant.
- Some points are under review
- 24 cases are waiting for a manual review: transcriptions with tone marks we could not confirm, municipality names that are not in INEGI's list, and two geographic misprints in the catalog. In the meantime, 13 of the 476 phonetic transcriptions are not published. None of this was filled in with a guess.
- One threshold in the risk grade is inferred
- The rules INALI publishes for assigning the grades reproduce 351 of the 364 published grades. A child-speaker threshold around 15% reproduces all 364 published grades under the stated rules for grade 2 vs 3; INALI's text does not state this value explicitly — it's a reverse-engineered inference from the data, confirmed against every known case but not confirmed as the source's actual rule. The site does not use it: it always shows the published grade.
- One column in the risk book is not explained
- The book publishes a "proportion of speakers" for each variant without saying which population it is a share of. We show it as published, without interpreting it.
- "Focused on the language" is our own criterion
- Many Hugging Face repositories tag hundreds of languages at once. To set those apart, we count separately the ones tagged for three groups or fewer. We chose that cutoff; a different one would change the figures.
License
The code is free to use under the MIT license. The data Glototeca produces (the table of correspondences, the weekly snapshots and the files derived from them) is published under Creative Commons Attribution 4.0 (CC BY 4.0): it may be used and adapted with attribution.
That license covers our work, not anyone else's. INEGI's figures, INALI's catalogs and Glottolog's classification belong to their authors and keep their own terms; Glototeca cannot relicense them. Anyone reusing that data should consult and cite the original source.
One case is still open: the risk grades were extracted from a 2012 INALI book that reserves all rights. They are not published under CC BY 4.0; redistributing them depends on INALI's answer, and the current status is on the Data page.
How to cite
If you use this data in an article, a news report or research, please cite it like this.
Rangel, J. (2026). Glototeca: Digital resource map of Mexico's indigenous languages [Data set]. Retrieved October 5, 2026, from https://glototeca.com/
Rangel, Jhonnatan. “Glototeca.” Glototeca, 5 Oct. 2026, https://glototeca.com/.
The data reflects the snapshot of October 5, 2026.Access date: October 5, 2026
There is no permanent DOI yet. Until there is one, cite the site address and your access date.
Indigenous data
This site records public metadata and links only: that something exists, where it is published and when it was seen. It does not copy, host or redistribute recordings, texts or materials in any language. Community-hosted and private content is never counted.
A zero in these tables does not say that a language lacks resources. It says nothing is published in the sources we check. The knowledge a community keeps to itself is not here, and should not be.
We take the CARE Principles for Indigenous Data Governance as our guide. Declaring them is not the same as meeting them; this is what we do today under each one:
- Collective benefit
- The site is open and free, meant for communities themselves to be able to see where their language stands.
- Authority to control
- We do not gather community data or make decisions about it. We only point to what others have already published.
- Responsibility
- Every figure carries its source and date, and the errors we find are flagged, not hidden.
- Ethics
- We do not rate languages or communities with a score, and we do not present missing data as a deficiency.