Open data · updated weekly

Which languages have data, models and corpora, and which do not?Which languages have data, models and corpora, and which do not?

Every week we record which datasets, models, speech corpora, treebanks and tools exist for Mexico's 68 language groups, set beside their risk grade and number of speakers.

Latest snapshot
Language groups
68
INALI's catalog
Language variants
364
Each with an INALI risk grade
Languages in the reference list
7,674
Glottolog 5.3, spoken L1 languages
Snapshot date
Oct 5, 2026
Weekly, append-only

Every value is stored with the date it was measured. Indicators are never blended into a single score.

More speakers does not always mean more resourcesMore speakers does not always mean more resources

Each dot is a language group. Further right, more speakers; further up, more kinds of digital resource. Color shows INALI's risk grade.

This means the technical building blocks for language tools — translators, speech recognition, spell-checkers — not cultural content like books or music. What we count and where it comes from

Risk grade (INALI 2012), highest-risk variant
  • Very high
  • High
  • Medium
  • No immediate risk

Kinds of resource with at least one record (of 5)

Speakers aged 3 and over (2020 Census, log scale)

The five kinds: datasets and models on Hugging Face focused on the language, a speech corpus in Common Voice, treebanks in Universal Dependencies, and support in Omnilingual ASR.

Within each row the dots are spread slightly up and down so they do not overlap.

Not shown in the chart: ku'ahl (no row in the census; identity unresolved).