Which languages have data, models and corpora, and which do not?Which languages have data, models and corpora, and which do not?
Every week we record which datasets, models, speech corpora, treebanks and tools exist for Mexico's 68 language groups, set beside their risk grade and number of speakers.
- Language groups
- 68
- INALI's catalog
- Language variants
- 364
- Each with an INALI risk grade
- Languages in the reference list
- 7,674
- Glottolog 5.3, spoken L1 languages
- Snapshot date
- Oct 5, 2026
- Weekly, append-only
Every value is stored with the date it was measured. Indicators are never blended into a single score.
More speakers does not always mean more resourcesMore speakers does not always mean more resources
Each dot is a language group. Further right, more speakers; further up, more kinds of digital resource. Color shows INALI's risk grade.
This means the technical building blocks for language tools — translators, speech recognition, spell-checkers — not cultural content like books or music. What we count and where it comes from
- Very high
- High
- Medium
- No immediate risk
Kinds of resource with at least one record (of 5)
Speakers aged 3 and over (2020 Census, log scale)
The five kinds: datasets and models on Hugging Face focused on the language, a speech corpus in Common Voice, treebanks in Universal Dependencies, and support in Omnilingual ASR.
Within each row the dots are spread slightly up and down so they do not overlap.
Not shown in the chart: ku'ahl (no row in the census; identity unresolved).