Method

What Glototeca measures, where the data comes from, how we decide that two names are the same language, and what we do not know yet.

What this measures and what it does not

Once a week, Glototeca records which public digital resources exist for each of the 68 language groups in INALI's catalog: datasets, models, speech corpora, treebanks and support in one speech-recognition tool. It sets them beside two pieces of context: how many people speak the language and how much at risk it is.

It measures presence in public sources, not quality or usefulness. A dataset existing for a language says nothing about whether it is good, whether the community knows about it or whether it serves them. And nothing showing up does not mean nothing exists: it means nothing shows up in the sources we check.

The four states of a value

3
Measured

Ayapaneco: 3 datasets on Hugging Face carry its tag. →

0
Checked, nothing found

Ayapaneco: we checked Universal Dependencies, which could hold a treebank for any language, and there is none. This is a real zero. →

Source does not cover this language
Source does not cover this language

Ayapaneco: Common Voice publishes a closed list of languages and this one is not on it. That is not a zero: the source does not include it. →

Identity unresolved
Identity unresolved

Ku’ahl: it has no ISO 639-3 code, so there is nothing to look it up by in any source. The value is unknown, and is shown as unknown. →

No single score

The indicators are never combined into a rating. Adding datasets to treebanks, or weighting one against another, would produce a number that looks precise and means nothing. Each indicator is shown separately, with its source and its date.

Mexico is the pilot. The long-term aim is to apply the same method to every language in Glottolog's reference list (7,674 spoken languages in version 5.3); today the site covers Mexico only.

Where the data comes from

Eight sources, all public and free. Five are checked every week; three are fixed publications that were loaded once. The latest snapshot is from Oct 5, 2026.

SourceWhat it providesHow it is obtainedCadenceVersion on record
GlottologEach language's identity and its endangerment statusA dated release, downloaded from Zenodo and pinned by its DOIWeeklyglottolog 5.3 (doi:10.5281/zenodo.18840967)
Hugging FaceDatasets and models tagged with the languagePublic API, no accountWeeklyOct 5, 2026
Common VoiceWhether a speech corpus exists, and for how many varietiesMetadata for each release, published on GitHubWeeklycommon-voice scripted 27.0 (2026-09-11), spontaneous 5.0 (2026-09-11)
Universal DependenciesReleased treebanksTwice-yearly releases on GitHubWeeklyuniversal-dependencies 2.18 (2026-05-15)
Omnilingual ASRWhether the speech-recognition model declares supportThe model's own language list, by codeWeeklyomnilingual-asr lang_ids.py @ a7fb36017a (2025-11-10)
INEGI, 2020 CensusSpeakers aged 3 and over, by languageOfficial census table, downloaded onceFixedINEGI Censo de Población y Vivienda 2020, cpv2020_b_eum_05_etnicidad.xlsx sheet 03
INALI, Catálogo de las Lenguas Indígenas Nacionales (2008)The 364 variants: names, autonyms and localitiesOfficial PDF, read by a program and verified by its checksumFixedPDF
INALI, Lenguas indígenas nacionales en riesgo de desaparición (2012)The risk grade of each variantOfficial PDF, read by a program and verified by its checksumFixedPDF

No source requires payment or a key. What a source publishes under its own terms stays under those terms.

How identity is resolved

Every source names languages its own way. INALI speaks of groups and variants; Hugging Face and Common Voice use ISO 639-3 codes; Glottolog uses its own identifiers. Before counting anything, someone has to decide which code belongs to which group. That table of correspondences was built once, by hand, and checked against Glottolog's own cross-references. Everything else reads it automatically.

Of the 68 groups, 65 are resolved: 43 correspond to a single Glottolog language, 21 bring several together and 1 is a macrolanguage. The remaining 3 (ku'ahl, otomí, zoque) are not resolved, and the site shows them that way. A doubtful identity is flagged; it is never guessed.

An error we found

The 2008 catalog and the 2012 risk book are two INALI publications about the same 364 variants. When we joined them, 354 names matched letter for letter. Among those that did not, 3 differ by a single word, and that word is a direction: where one publication says "noreste" (northeast), the other says "noroeste" (northwest).

This is not a spelling variant. They are two different directions. What shows that each pair is one and the same variant is a different kind of evidence: the number of localities each publication assigns to it.

2008 catalog2012 risk bookLocalities (2008)Localities (2012)
zapoteco de la Sierra sur, noroestezapoteco de la Sierra sur, noreste1616
zapoteco de la Sierra sur, noreste altozapoteco de la Sierra sur, noroeste alto1313
náhuatl del noreste centralnáhuatl del noroeste central400398

The error is most likely in the 2008 catalog, though we cannot rule out that it is in the 2012 book. The documents alone do not settle it.

So we did not correct it. The site keeps the name exactly as the catalog prints it, links each variant to its row in the risk book, and says on the variant's page how the link was made. The discrepancy stays flagged in the data, in plain view, so that INALI or any specialist can resolve it.

Limitations

What to know before quoting a figure from this site.

Risk and speakers come from different censuses
The risk grades were calculated from the 2000 census. The speaker counts we show are from the 2020 census. Twenty years separate them: a variant may be better or worse off today than its grade says.
Most variants have no identifier of their own
Only 136 of the 364 variants carry a Glottocode, and it is their group's, not their own. The other 228 belong to groups that correspond to several Glottolog languages, and there is no reliable way to assign them. That is why digital resources are measured per group, not per variant.
Some points are under review
24 cases are waiting for a manual review: transcriptions with tone marks we could not confirm, municipality names that are not in INEGI's list, and two geographic misprints in the catalog. In the meantime, 13 of the 476 phonetic transcriptions are not published. None of this was filled in with a guess.
One threshold in the risk grade is inferred
The rules INALI publishes for assigning the grades reproduce 351 of the 364 published grades. A child-speaker threshold around 15% reproduces all 364 published grades under the stated rules for grade 2 vs 3; INALI's text does not state this value explicitly — it's a reverse-engineered inference from the data, confirmed against every known case but not confirmed as the source's actual rule. The site does not use it: it always shows the published grade.
One column in the risk book is not explained
The book publishes a "proportion of speakers" for each variant without saying which population it is a share of. We show it as published, without interpreting it.
"Focused on the language" is our own criterion
Many Hugging Face repositories tag hundreds of languages at once. To set those apart, we count separately the ones tagged for three groups or fewer. We chose that cutoff; a different one would change the figures.

License

The code is free to use under the MIT license. The data Glototeca produces (the table of correspondences, the weekly snapshots and the files derived from them) is published under Creative Commons Attribution 4.0 (CC BY 4.0): it may be used and adapted with attribution.

That license covers our work, not anyone else's. INEGI's figures, INALI's catalogs and Glottolog's classification belong to their authors and keep their own terms; Glototeca cannot relicense them. Anyone reusing that data should consult and cite the original source.

One case is still open: the risk grades were extracted from a 2012 INALI book that reserves all rights. They are not published under CC BY 4.0; redistributing them depends on INALI's answer, and the current status is on the Data page.

How to cite

If you use this data in an article, a news report or research, please cite it like this.

APA 7

Rangel, J. (2026). Glototeca: Digital resource map of Mexico's indigenous languages [Data set]. Retrieved October 5, 2026, from https://glototeca.com/

MLA 9

Rangel, Jhonnatan. “Glototeca.” Glototeca, 5 Oct. 2026, https://glototeca.com/.

The data reflects the snapshot of October 5, 2026.Access date: October 5, 2026

There is no permanent DOI yet. Until there is one, cite the site address and your access date.

Indigenous data

This site records public metadata and links only: that something exists, where it is published and when it was seen. It does not copy, host or redistribute recordings, texts or materials in any language. Community-hosted and private content is never counted.

A zero in these tables does not say that a language lacks resources. It says nothing is published in the sources we check. The knowledge a community keeps to itself is not here, and should not be.

We take the CARE Principles for Indigenous Data Governance as our guide. Declaring them is not the same as meeting them; this is what we do today under each one:

Collective benefit
The site is open and free, meant for communities themselves to be able to see where their language stands.
Authority to control
We do not gather community data or make decisions about it. We only point to what others have already published.
Responsibility
Every figure carries its source and date, and the errors we find are flagged, not hidden.
Ethics
We do not rate languages or communities with a score, and we do not present missing data as a deficiency.