Data

What is measured, where each figure comes from, and the terms its source publishes it under. Glototeca's own downloads are not available yet; every original source is, and each one is linked here.

What we mean by digital resources

This means the technical building blocks for language tools: translators, speech recognition, spell-checkers, predictive keyboards. It does not mean cultural content such as books, music, video or websites written in the language.

The distinction matters because they are different things. A language can have a large literature and no corpus a model could be trained on; the reverse also happens. This site measures only the second, and says nothing about a language's cultural vitality.

The five kinds that are counted

Datasets
Collections of text or audio prepared for a program to read.
Models
Trained systems that declare they work with the language.
Speech corpora
Recordings with their transcriptions, the basis of speech recognition.
Treebanks
Sentences with their grammatical analysis annotated by hand.
Speech recognition support
Whether a recognition model lists the language among those it supports.

How each one is measured, and the four states a value can have, are on the Method page.

The sources, one by one

Eight sources. For each: what it provides, how Glototeca reads it, the terms it is published under, and where to get the original.

Glottolog

Weekly

Gives each language a stable identifier and its endangerment status; it is how every other source is tied to the right language.

How it is obtained
Dated release pinned by its DOI, downloaded from Zenodo
Cadence
Weekly: scheduled fetch every Monday at 06:17 UTC
Recorded version
glottolog 5.3 (doi:10.5281/zenodo.18840967)
Last measured
Source licence or terms
CC BY 4.0 (checked Oct 5, 2026)

Stated on the Zenodo record for release 5.3.

Original source
Go to the sourcedoi.org/10.5281/zenodo.18840967
Glototeca's processed version
Coming soon The download does not exist yet.

Hugging Face

Weekly

Counts the datasets and models tagged with the language in the largest public repository of both.

How it is obtained
Public API, no account or key
Cadence
Weekly: scheduled fetch every Monday at 06:17 UTC
Recorded version
huggingface
Last measured
Source licence or terms
Per repository, plus the platform Terms of Service (checked Oct 5, 2026)

Each dataset and model states its own licence. Glototeca only counts the listings; it does not copy their content.

Original source
Go to the sourcehuggingface.co/docs/hub/api
Glototeca's processed version
Coming soon The download does not exist yet.

Common Voice

Weekly

Shows whether an open speech corpus exists for the language, the raw material of speech recognition.

How it is obtained
Files published in the source's public repository
Cadence
Weekly: scheduled fetch every Monday at 06:17 UTC
Recorded version
common-voice scripted 27.0 (2026-09-11), spontaneous 5.0 (2026-09-11)
Last measured
Source licence or terms
CC0-1.0 on Scripted Speech 27.0 (checked Oct 5, 2026)

Unconfirmed for Spontaneous Speech 5.0: its listing with the licence was not found. The metadata repository that is read is MPL-2.0.

Original source
Go to the sourcegithub.com/common-voice/cv-dataset
Glototeca's processed version
Coming soon The download does not exist yet.

Universal Dependencies

Weekly

Counts released treebanks, the resource grammatical parsers are built from.

How it is obtained
Files published in the source's public repository
Cadence
Weekly: scheduled fetch every Monday at 06:17 UTC
Recorded version
universal-dependencies 2.18 (2026-05-15)
Last measured
Source licence or terms
Per treebank (checked Oct 5, 2026)

Each treebank carries its own licence, usually a Creative Commons variant; some do not allow commercial use.

Original source
Go to the sourceuniversaldependencies.org/
Glototeca's processed version
Coming soon The download does not exist yet.

Omnilingual ASR

Weekly

Shows whether a wide-coverage speech recognition model declares support for the language.

How it is obtained
Files published in the source's public repository
Cadence
Weekly: scheduled fetch every Monday at 06:17 UTC
Recorded version
omnilingual-asr lang_ids.py @ a7fb36017a (2025-11-10)
Last measured
Source licence or terms
Apache 2.0 (code and models); CC-BY-4.0 (its corpus) (checked Oct 5, 2026)
Original source
Go to the sourcegithub.com/facebookresearch/omnilingual-asr
Glototeca's processed version
Coming soon The download does not exist yet.

INEGI, 2020 Population and Housing Census

Fixed

Gives the number of speakers aged 3 and over per language, against which resources are read.

How it is obtained
Official table, downloaded directly
Cadence
Fixed: a publication loaded once, not re-measured
Recorded version
INEGI Censo de Población y Vivienda 2020, cpv2020_b_eum_05_etnicidad.xlsx sheet 03
Source licence or terms
INEGI Terms of Free Use (Términos de libre uso) (checked Oct 5, 2026)

They allow copying, distributing, adapting and commercial use. They require credit as "Fuente: INEGI, [product]" and disclosure of any transformation, without presenting it as done or endorsed by INEGI.

Original source
Go to the sourcewww.inegi.org.mx/programas/ccpv/2020/#tabulados
Glototeca's processed version
Coming soon The download does not exist yet.

INALI, Catálogo de las Lenguas Indígenas Nacionales (2008)

Fixed

Defines the 68 groups and their 364 variants, with autonyms and localities: the site's unit of record.

How it is obtained
Official PDF, read by a program and verified by its checksum
Cadence
Fixed: a publication loaded once, not re-measured
Recorded version
https://www.inali.gob.mx/pdf/CLIN_completo.pdf
Source licence or terms
No license stated (checked Oct 5, 2026)

The document, published in the Diario Oficial de la Federación on 14 January 2008, contains no rights notice, licence or reproduction terms. All 256 pages were checked.

Original source
Go to the sourcewww.inali.gob.mx/pdf/CLIN_completo.pdf
Glototeca's processed version
Coming soon The download does not exist yet.

INALI, Lenguas indígenas nacionales en riesgo de desaparición (2012)

Fixed

Gives each variant's risk grade, shown beside every resource measurement.

How it is obtained
Official PDF, read by a program and verified by its checksum
Cadence
Fixed: a publication loaded once, not re-measured
Recorded version
https://site.inali.gob.mx/pdf/libro_lenguas_indigenas_nacionales_en_riesgo_de_desaparicion.pdf
Source licence or terms
All rights reserved (D.R. © 2012 INALI) (checked Oct 5, 2026)

The legal page prohibits reproduction in whole or in part without written permission.

The risk grades shown on this site were extracted from this all-rights-reserved publication. Permission to redistribute them as open data, with attribution, will be requested from INALI; it has not yet been asked for or confirmed.

Original source
Go to the sourcesite.inali.gob.mx/pdf/libro_lenguas_indigenas_nacionales_en_riesgo_de_desaparicion.pdf
Glototeca's processed version
Coming soon The download does not exist yet.

A note on licences

Two things with different terms sit side by side here.

Glototeca's own work is published under CC BY 4.0: the crosswalk between groups, ISO 639-3 codes and Glottocodes; the weekly measurements with their four states; the variant-level resolution; and the findings on direction errors in variant names. It will be downloadable from this page once that feature exists.

Data drawn from third parties keeps whatever licence each source publishes, shown above with the date it was checked. Glototeca does not claim the right to redistribute it in bulk until that is confirmed source by source.

One case is open. The risk grades come from a 2012 INALI book that reserves all rights. Their status is not settled either way: they are not published under CC BY 4.0 as if it were, and they depend on INALI's answer. The current status is on that source's card.

Formal licence statement, on the Method page