Human Tech Tree
Current research2022-2026 · Present (2015 – now)

Civilization / Humanities & Culture

AI for Low-Resource Languages

Meta's NLLB (200 languages), Omnilingual ASR (1,600+ languages) and 110 languages added to Google Translate in 2024 reach languages that older systems ignored.

Open in the interactive tree →

Models trained on many languages at once pass what they learn from large languages to small ones. Meta's NLLB project claimed a 44% BLEU gain over the previous best on a human-translated 200-language test set (2022 preprint) and released models and data openly. Quality still falls with the amount of real text and speech available for a language.

As of now

Omnilingual ASR (Meta, 10 November 2025) transcribes speech in 1,600+ languages, 500 never covered by AI before, yet only 36% of languages with under ten hours of audio reach a character error rate below 10%. Google added 110 languages to Translate on 27 June 2024, for 614 million speakers. A June 2026 preprint found human scores of 4.0-4.5 out of 5 for Hausa but 1.0-2.2 for Fongbe from the same models, and a 2024 Mambai study scored BLEU 21.2 on language-manual text but 4.4 on a native speaker's sentences. Ethnologue counts over 7,000 living languages (2025).

Open steps

  • Data for the smallest languages Medium AI leverageFind text and speech for languages with under ten hours of audio, where only 36% reach under 10% character error in Meta's Omnilingual ASR.
  • Tests native speakers trust Medium AI leverageBuild test sets written by native speakers; a Mambai study scored BLEU 21.2 on language-manual text but 4.4 on a native speaker's sentences.
  • Who controls the data Low AI leverageAgree who owns and may reuse community recordings; Te Hiku Media's Kaitiakitanga Licence limits use of its Māori speech data to the benefit of Māori people.
  • Languages without writing Medium AI leverageTurn speech into text or other languages for oral languages; Omnilingual ASR transcribes 1,600+ languages, but transcription is not translation.

Where AI could help

Medium AI leverage. The models already do the work, but quality depends on text and speech that small language communities hold, not on more compute alone.

  • Transfer knowledge from large languages to small ones
  • Turn speech into text so spoken languages gain written data
  • Mine aligned sentences from existing texts and recordings
  • Build keyboards, spell-checkers and speech tools for small languages

Shown so far

  • Meta's NLLB project reports a 44% BLEU improvement over the previous state of the art across 200 languages and open-sources its models (2022 preprint). source
  • Meta's Omnilingual ASR (10 November 2025) covers 1,600+ languages; 78% reach under 10% character error, but only 36% of languages with under ten hours of audio do. source
  • On 27 June 2024 Google added 110 languages to Translate, spoken by 614 million people, using its PaLM 2 model (company announcement). source

Prerequisites

Unlocks

Sources

More in Humanities & Culture · Present

All 41 points in Humanities & Culture →

Open in the interactive tree →