New Report on SA Gambling Impact
Check It Out
<-BackA language-by-language guide to AI transcription and translation for research: 130 languages, 23 markets, which provider wins where, and when you still need a human.

AI Transcription & Translation for Research: A Language Guide

Data Analysis
Created at:
July 9, 2026
Updated at:
October 8, 2026
Field Guide · Multilingual Research · Updated October 2026

You have research to run in Swahili, or isiZulu, or Hindi, or across a dozen markets at once, and you want one answer: can AI transcribe and translate it, or do you still need people doing it by hand? This is our language-by-language answer, rebuilt in October 2026 with 130 languages, 23 markets and a head-to-head of the main providers.

Topic
AI Transcription
Languages rated
130
Read time
25 minutes
Updated
October 2026
130
Languages rated for AI transcription and translation, up from 74 in August.
23 of 74
August ratings that moved when we checked them against current providers and benchmarks.
6% vs 34%
Hindi words transcribed wrong on the same phone calls by the best and worst big provider.

The short version: for English, Spanish and the big European and Asian languages, AI now does the work and a light check is enough. For most of Africa's large languages, AI gives you a good draft that a native speaker needs to check. For a long tail, it still can't hear the language at all. And two things cut across every rating: which provider you use, and who is speaking. The same set of Hindi phone calls can come back with 6% of words wrong or 34%, depending on the engine. The same system can get English near perfect when it's read into a studio microphone, and a third wrong on calls to an advice charity in North East Scotland.

What changed since August

We published the first version of this guide in July and updated it in August. When we checked it again in October, 23 of the 74 ratings had moved, and we'd missed a lot of languages people actually field in. So we rebuilt it against what the providers support today and the newest real-world benchmarks.

  • Up. Cantonese, Punjabi, Odia, Assamese and Nepali are now Ready, and so is Levantine Arabic, though that rests on one benchmark. Bhojpuri, Luganda, Wolof and Nigerian Pidgin moved from Hard to Workable. Pidgin finally has a commercial speech model (Intron Sahara v2, March 2026). Translation into isiZulu, Kinyarwanda and Javanese is now Strong, and Xitsonga, siSwati and Kirundi moved up to Moderate.
  • Down. Hausa and Egyptian Arabic dropped to Workable once we looked at real recordings rather than clean test audio. Moroccan Darija, Sudanese Arabic and Kurmanji Kurdish dropped to Hard. Hebrew is now a borderline Workable. Translation into Wolof and Fula dropped to Basic.
  • New. 56 languages added, including Turkish, Malay, Gulf and Iraqi Arabic, Pashto, Dari, Central Asian languages, 14 more sub-Saharan African languages, Tamazight and Kabyle, Haitian Creole, and indigenous languages of Latin America such as Quechua, Aymara, Guarani, Nahuatl and the main Mayan languages.
  • Advice we've changed. We used to suggest transcribing first and letting an LLM clean up mixed-language text. A 2026 study found that made Hindi transcripts worse. We also used to point to Meta's MMS model for the long tail. Meta has since replaced it with Omnilingual ASR (November 2025, open source, 1,600+ languages), though it still leaves out Sesotho, siSwati, Tshivenda and isiNdebele.
  • New sections. Which provider wins where, which model translates best, how accents change the picture inside “ready” languages, and the language mix inside 23 markets.

How the ratings work

Two separate ratings, because they're two separate jobs. Transcription turns a voice note into text in the same language. Translation turns that text into another language. Each language is rated on the best tool available on the market, not on any one provider, and we apply the same rules to every language.

  • ●
    Ready (transcription). At least two commercial providers, and either ElevenLabs rates its own accuracy at 10% word error or better, or the best real-world benchmark is around 15% or better.
  • ◐
    Workable. At least one current commercial provider, including regional specialists such as Intron, Sarvam and Lelapa, with real-world error of roughly 15 to 35%, or a listing with no published accuracy.
  • ○
    Hard. No current commercial speech-to-text, or best evidence worse than about 35 to 40% of words wrong.
  • ●
    Strong / Moderate / Basic (translation). Based on how closely frontier models (GPT, Claude and Gemini) match a reference translation into the language, on the AI Language Proficiency Monitor's 0 to 100 chrF score: Strong at 57 or above, Moderate from about 42 to 56 or listed by Google Translate or DeepL with no benchmark, Basic below that. We rate up where the score is known to undercount, such as Ge'ez script (Amharic, Tigrinya) and Yoruba tone marks.
ASR (automatic speech recognition) is the technical name for speech-to-text. Word error rate is the share of words a system gets wrong, so lower is better. For Chinese, Japanese and Korean, benchmarks count characters instead (character error rate), which can't be compared directly with a word error rate. Real phone and voice-note audio usually scores 1.5 to 2 times worse than the clean recordings providers quote.

The language table

All 130 languages, grouped by region. The last column tells you what to do in fieldwork and the evidence behind the rating. Moved marks a rating that changed since August, New a language we added. Speaker counts are broad public estimates.

● Ready AI does the job, light spot-check.  ◐ Workable AI drafts, a native speaker checks.  ○ Hard Human-led, AI assists.

LanguageSpeakersTranscriptionTranslationWhat to do, and the evidence
Southern Africa
Afrikaans12M● Ready● StrongAI, light spot-check.ElevenLabs Good (10-20%); open-model bench 9.6%; chrF 68
isiZuluMoved28M◐ Workable● StrongAI drafts, a native speaker checks.Translation up: chrF 61, top-ranked LLMs in WMT24++. ASR 16-26% (AfriVox), 48%+ code-switched
isiXhosa19M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; AfriVox 17-29%; chrF 52
Sesotho14M◐ Workable◐ ModerateAI drafts, a native speaker checks.Commercial via Google legacy, Intron, Lelapa; AfriVox best 15%; not in Meta Omnilingual
Sepedi14M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; Chirp 3 preview (nso-ZA); chrF 53 (thin evidence)
Setswana14M◐ Workable◐ ModerateAI drafts, a native speaker checks.Google legacy + Intron; AfriVox best 11.6% (vendor bench); chrF 48
Shona15M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; AfriVox 21-37%
Chichewa20M◐ Workable◐ ModerateAI drafts, a native speaker checks.Google chirp_2 only; best 22%, most results 30%+ (thin evidence)
XitsongaMoved3.5M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Translation up: chrF 56, Google Translate 63. ASR: Google legacy model only
siSwatiMoved3M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Translation up: Google Translate added 2024, chrF 51. ASR: no modern commercial model
Tshivenda2M○ Hard○ BasicHuman-led, AI assists.Google legacy ASR only; translation consumer-app only, no benchmark (thin evidence)
isiNdebele1.5M○ Hard○ BasicHuman-led, AI assists.No commercial ASR; open models only (thin evidence)
East Africa
Swahili95M● Ready● StrongAI, light spot-check.ElevenLabs High (5-10%); AfriVox 7%; chrF 70
Amharic78M◐ Workable◐ ModerateAI drafts, a native speaker checks.AfriVox ~25%; chrF 37 into / 55 out of Amharic (script undercounts)
Somali22M◐ Workable◐ ModerateAI drafts, a native speaker checks.AWS, Azure, Google; ElevenLabs Moderate; open bench 35% (thin evidence)
KinyarwandaMoved15M◐ Workable● StrongAI drafts, a native speaker checks.Translation up: chrF 60. ASR mixed: 6.6% (Intron, own bench) to 30%+ elsewhere
LugandaMoved10M◐ Workable◐ ModerateAI drafts, a native speaker checks.Transcription up: AWS, Google, ElevenLabs, Intron now commercial; best 16% (Sunbird)
Oromo46M◐ Workable◐ ModerateAI drafts, a native speaker checks.Google chirp_2, Intron; open bench 21-25%; chrF 43
Tigrinya9M○ Hard○ BasicHuman-led, AI assists.No commercial ASR; chrF 30 into Tigrinya
KirundiMoved12M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Translation up: Google Translate added 2024, chrF 49. No commercial ASR
DholuoNew◐ Workable◐ ModerateAI drafts, a native speaker checks.Google (luo-KE) and ElevenLabs list it; open bench 17.5% (thin evidence)
KikuyuNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR (open bench 22.6%); not in Google Translate
MalagasyNew◐ Workable● StrongAI drafts, a native speaker checks.AssemblyAI and Gladia list it; open bench 13.7%; chrF 57 (thin evidence)
DinkaNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; Google Translate 2024 (thin evidence)
West & Central Africa
HausaMoved94M◐ Workable● StrongAI drafts, a native speaker checks.Transcription down: ElevenLabs Good (10-20%); best real-world 17-27%; 32% code-switched
Yoruba50M◐ Workable◐ ModerateAI drafts, a native speaker checks.AfriVox 27-36%; 45%+ code-switched; translation drops tone marks (chrF 33)
Igbo30M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; best 24%; chrF 53
Lingala40M◐ Workable◐ ModerateAI drafts, a native speaker checks.Google chirp_2, AssemblyAI; open bench 21.5% (thin evidence)
Akan (Twi)20M◐ Workable◐ ModerateAI drafts, a native speaker checks.Intron; best 28.7%; chrF 46
WolofMoved12M◐ Workable○ BasicHuman-led, AI assists.Transcription up (AWS, Chirp 3 preview, Intron; best ~28-32%). Translation down: chrF 34
FulaMoved40M○ Hard○ BasicHuman-led, AI assists.Translation down: chrF 25. ASR best 34%, most results 40-70%
Nigerian PidginMoved60-120M◐ Workable◐ ModerateAI drafts, a native speaker checks.Transcription up: Intron Sahara v2 supports it (Mar 2026); open bench 19.7%. Not in Google/Azure (thin evidence)
Mossi9M○ Hard○ BasicHuman-led, AI assists.No commercial ASR; not in Google Translate; chrF 29
Kanuri9M○ Hard○ BasicHuman-led, AI assists.No commercial ASR; open bench 35%; Google Translate app only (thin evidence)
EweNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No commercial ASR (open bench 29%); Google Translate (thin evidence)
BambaraNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; open bench 44%; chrF 34
FonNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No commercial ASR; Google Translate 2024 (thin evidence)
KrioNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No commercial ASR; English-based creole, Google Translate (thin evidence)
SangoNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No commercial ASR; Google Translate 2024 (thin evidence)
BembaNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Intron lists it (no figure); open bench 37-44%; chrF 48
TshilubaNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; NLLB only (thin evidence)
MakhuwaNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; no Google Translate (thin evidence)
UmbunduNew◐ Workable○ BasicHuman-led, AI assists.Listed by ElevenLabs and Google with no accuracy figure; not in Google Translate (thin evidence)
OshiwamboNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; no Google Translate (thin evidence)
Middle East & North Africa
Arabic (MSA)335M● Ready● StrongAI, light spot-check.Broad coverage; ElevenLabs Good; best open model 10.2% on MSA
Egyptian ArabicMoved118M◐ Workable● StrongAI drafts, a native speaker checks.Transcription down: best real-world 37% (Alibaba, GigaSpeechBench); best Western provider Gemini 41%, ElevenLabs 44%
Levantine ArabicMoved58M● Ready◐ ModerateAI drafts, a native speaker checks.Transcription up: real-world Syria best 13.8% (Alibaba), Gemini 14.4%, ElevenLabs 14.7% (thin evidence)
Gulf ArabicNew◐ Workable◐ ModerateAI drafts, a native speaker checks.Google, Azure, AWS (ar-AE), Deepgram; real-world Saudi best 16.6% (Alibaba), Google Chirp 3 16.8%; UAE 26%
Iraqi ArabicNew◐ Workable◐ ModerateAI drafts, a native speaker checks.Azure, Deepgram, Google ar-IQ; real-world best 28.5% (Alibaba), Microsoft Azure 34.6%
Moroccan DarijaMoved21M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Transcription down: best real-world 51%, every system above 44%
Algerian & Tunisian ArabicNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Locales exist (Azure, Google) but best real-world 44% (thin evidence)
Sudanese ArabicMoved54M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Transcription down: Deepgram ar-SD only; best open model 36% (thin evidence)
TamazightNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; Google Translate 2024 (Latin and Tifinagh) (thin evidence)
KabyleNew◐ Workable◐ ModerateAI drafts, a native speaker checks.AWS Transcribe (kab-DZ), no accuracy published; open 29% (thin evidence)
Persian82M● Ready● StrongAI, light spot-check.ElevenLabs High (5-10%); broad coverage
Kurdish (Kurmanji)Moved26M○ Hard◐ ModerateAsk for typed answers, or a human transcribes.Transcription down: Kurmanji not supported by any major provider (Sorani only)
Kurdish (Sorani)New◐ Workable◐ ModerateAI drafts, a native speaker checks.AWS and Google (ckb-IQ, ckb-IR); ElevenLabs 'Kurdish' Moderate (thin evidence)
HebrewMoved9M◐ Workable● StrongAI drafts, a native speaker checks.Transcription down to the rule: ElevenLabs Good (10-20%), no real-world benchmark. Borderline (thin evidence)
TurkishNew● Ready● StrongAI, light spot-check.ElevenLabs Excellent (5% or under); every major provider
South Asia
Hindi611M● Ready● StrongAI, light spot-check.Real-world 5-8% (Voice of India)
Bengali274M● Ready● StrongAI, light spot-check.Real-world 6-10%
Urdu246M● Ready● StrongAI, light spot-check.Indian phone audio: Sarvam 7.0%, Gemini 3 Pro 9.1%. ElevenLabs rates it Moderate (25 to 50%); Google Chirp 3 has no Pakistani Urdu locale
Tamil86M● Ready● StrongAI, light spot-check.ElevenLabs High; real-world best 14%
Telugu96M● Ready● StrongAI, light spot-check.ElevenLabs High; real-world best 18%
Marathi99M● Ready● StrongAI, light spot-check.Real-world 9-11%
Gujarati62M● Ready● StrongAI, light spot-check.Real-world best 12.8%
Kannada59M● Ready● StrongAI, light spot-check.ElevenLabs Excellent; real-world best 16%
Malayalam38M● Ready● StrongAI, light spot-check.ElevenLabs Excellent; real-world best 19%
PunjabiMoved90M● Ready● StrongAI, light spot-check.Transcription up: Indian phone audio best 11.2% (Sarvam), Gemini 14.4%. Pakistani Punjabi (Shahmukhi) not benchmarked; treat as Workable in Pakistan
OdiaMoved38M● Ready◐ ModerateAI drafts, a native speaker checks.Transcription up: ElevenLabs High; real-world best 14%
AssameseMoved24M● Ready◐ ModerateAI drafts, a native speaker checks.Transcription up: real-world best 12.7%
NepaliMoved32M● Ready● StrongAI, light spot-check.Transcription up: ElevenLabs High (5-10%)
Sindhi37M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; Sarvam covers it
Sinhala26M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; Azure, AWS
BhojpuriMoved53M◐ Workable◐ ModerateAI drafts, a native speaker checks.Transcription up: real-world best 18.4% (Gemini 3 Pro), Sarvam 20.9%
MaithiliNew◐ Workable◐ ModerateAI drafts, a native speaker checks.Sarvam covers it; real-world best 24.7%
PashtoNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate (25-50%); AWS, Azure, Deepgram, AssemblyAI
DariNew◐ Workable◐ ModerateAI drafts, a native speaker checks.AWS fa-AF; otherwise relies on Persian models (thin evidence)
SaraikiNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; not in Google Translate (thin evidence)
East & Southeast Asia
Mandarin1,184M● Ready● StrongAI, light spot-check.ElevenLabs High
Japanese126M● Ready● StrongAI, light spot-check.ElevenLabs Excellent on clean audio, but real-world 28 to 44% of characters wrong across Western providers (partly scoring: kanji vs kana)
Korean82M● Ready● StrongAI, light spot-check.Real-world best 9.9% of characters wrong (Alibaba FunASR, GigaSpeechBench); ElevenLabs 11.8%, Microsoft 13.1%
Indonesian255M● Ready● StrongAI, light spot-check.ElevenLabs Excellent; real-world best 15% (Alibaba), Google Chirp 3 20%
Vietnamese97M● Ready● StrongAI, light spot-check.Real-world best 9.6% (Google Chirp 3)
Thai71M● Ready● StrongAI, light spot-check.Real-world best 10.8% (Alibaba), ElevenLabs 13.9%
Tagalog87M● Ready● StrongAI, light spot-check.ElevenLabs High; with Taglish mixing, real-world best about a quarter of words wrong
MalayNew● Ready● StrongAI, light spot-check.ElevenLabs Excellent; every major provider
CantoneseMoved86M● Ready● StrongAI, light spot-check.Real-world best 6.1% CER (Alibaba FunASR); Microsoft Azure 11.8%, Google Chirp 3 48%
Cebuano28M◐ Workable● StrongAI drafts, a native speaker checks.Listed by ElevenLabs, Google, Gemini with no accuracy figure (thin evidence)
JavaneseMoved69M◐ Workable● StrongAI drafts, a native speaker checks.Translation up: chrF 60. ASR ElevenLabs Good
Burmese43M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Good; chrF 56
Khmer18M◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate
LaoNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; Google, Azure
SundaneseNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No major commercial ASR; Google, DeepL translate it (thin evidence)
Taiwanese HokkienNew○ Hard○ BasicHuman-led, AI assists.ElevenLabs 71%, Azure 67% real-world; specialist open models ~28-30% CER
Central Asia & Caucasus
AzerbaijaniNew● Ready◐ ModerateAI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure (thin evidence)
KazakhNew● Ready◐ ModerateAI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram. Russian mixing common
UzbekNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Good (10-20%); Google, AWS, Azure
KyrgyzNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Good; Google preview (thin evidence)
TajikNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Good; not Azure or Chirp 3 (thin evidence)
GeorgianNew● Ready◐ ModerateAI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram
ArmenianNew● Ready◐ ModerateAI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram
MongolianNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate (25-50%); Google, AWS, Azure
Americas & Caribbean
Haitian CreoleNew◐ Workable● StrongAI drafts, a native speaker checks.AWS Transcribe (ht-HT) and AssemblyAI list it, no accuracy published; chrF 57 (thin evidence)
QuechuaNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; translation weak (AmericasNLP)
GuaraniNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR; chrF 39
AymaraNew○ Hard○ BasicHuman-led, AI assists.No commercial ASR
NahuatlNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text; Google Translate app only (Eastern Huasteca)
Yucatec MayaNew○ Hard◐ ModerateAsk for typed answers, or a human transcribes.No commercial speech-to-text; Google Translate incl. Cloud API (thin evidence)
TseltalNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
TsotsilNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
MixtecNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
ZapotecNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text; Google Translate app only
K'iche'New○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
Q'eqchi'New○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text; Google Translate app only
MamNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text; Google Translate app only
KaqchikelNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
KichwaNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text; translation only via generic Quechua
MapudungunNew○ Hard○ BasicHuman-led, AI assists.No commercial speech-to-text or machine translation
Europe & global
English1.5B● Ready● StrongAI, light spot-check.Leaderboard 2 to 5% across all big providers; regional accents far higher (real Scottish speech, best system 24%)
Spanish561M● Ready● StrongAI, light spot-check.Choose the regional model (es-MX, es-US, es-ES). Spanglish roughly doubled errors in a 2024 study of open models
French334M● Ready● StrongAI, light spot-check.Québec French 8 to 10% in a 2025 test; standard French benchmarks didn't predict which model did best
Portuguese269M● Ready● StrongAI, light spot-check.Choose pt-BR or pt-PT: a model tuned on European Portuguese made twice the errors on Brazilian
Russian210M● Ready● StrongAI, light spot-check.
German133M● Ready● StrongAI, light spot-check.
Italian66M● Ready● StrongAI, light spot-check.
Other European● Ready● StrongAI, light spot-check.Serbian checked: ElevenLabs High
AlbanianNew◐ Workable◐ ModerateAI drafts, a native speaker checks.Azure only; not in ElevenLabs bands or Chirp 3 (thin evidence)
WelshNew◐ Workable● StrongAI drafts, a native speaker checks.ElevenLabs Good; Azure, Speechmatics
IrishNew◐ Workable◐ ModerateAI drafts, a native speaker checks.ElevenLabs Moderate; Azure, Speechmatics

About 36 of these ratings rest on thin evidence, mostly newer African languages where all we have is a provider listing or an open-model test. They're marked in the table. The ratings move quickly, so check the date on anything you rely on.

Languages inside each country

A language rating only goes so far, because nobody runs a study in “isiZulu”. They run it in South Africa, where English is spoken at home by fewer than one in ten people. So we mapped the home languages of 23 markets, using censuses and statistics-office surveys, and coloured each language by its transcription rating.

World map with the home-language mix of 23 markets, coloured by AI transcription rating
Share of people by home language, coloured by AI transcription rating. Sources: national censuses and statistics-office surveys 2007 to 2025; Afrobarometer Round 10 (2024) for Nigeria, Kenya, Ghana, Uganda and Tanzania.

The split is stark. In Brazil, the UK, the US, Mexico and Spain, more than nine in ten people speak a home language AI transcribes well. In Ethiopia, Uganda, Nigeria, Ghana and Pakistan, fewer than one in ten do. South Africa sits at about one in five, because English and Afrikaans are Ready but isiZulu, isiXhosa, Sepedi, Setswana and Sesotho are only Workable.

Market Home language rated Ready Largest home languages that are not Ready
Brazil >99% None above 1% of people
United Kingdom 96% None above 1% of people
United States 95% None above 1% of people
Mexico 94% Tseltal, Tsotsil, Mixtec, Zapotec 2%, Nahuatl 1%
Spain 93% None above 1% of people
Canada 90% None above 1% of people
Germany 88% None above 1% of people
Turkey 86% Kurdish/Zazaki 13%
Peru 83% Quechua 14%, Aymara 2%
Bolivia 78% Quechua 13%, Aymara 7%
India 76% Bhojpuri 4%, Maithili 1%
Tanzania 76% Sukuma 8% (not rated)
Guatemala 70% Q'eqchi' 8%, K'iche' 8%, Mam 4%
Nepal 46% Maithili 11%, Bhojpuri 6%, Tharu 6% (not rated)
Philippines 40% Bisaya/Cebuano 22%, Hiligaynon 7% (not rated), Ilocano 7% (not rated)
Kenya 33% Kikuyu 13%, Kalenjin 8% (not rated), Kamba 8% (not rated)
Paraguay 31% Guarani + Spanish 39% (both), Guarani 28%
South Africa 19% isiZulu 24%, isiXhosa 16%, Sepedi 10%
Pakistan 9% Punjabi 37%, Pashto 18%, Sindhi 14%
Ghana 7% Akan 52%, Ewe 12%, Dagbani 7% (not rated)
Nigeria 6% Hausa 32%, Yoruba 17%, Igbo 11%
Uganda 3% Luganda 28%, Runyankore 10% (not rated), Lusoga 10% (not rated)
Ethiopia 0% Oromo 34%, Amharic 29%, Somali 6%
Reading the table. Mexico counts people who speak an indigenous language, most of whom also speak Spanish. In Paraguay, many homes use both Guarani and Spanish. On the country map further down, Kenya, Tanzania and India are rated on the shared languages used in fieldwork (Swahili and Hindi), so they look better there than in this table. Punjabi is Ready on Indian audio, but Pakistani Punjabi is written in a different script (Shahmukhi) that no provider publishes accuracy for, so we treat it as Workable in Pakistan. Urdu stays Ready: spoken Urdu is much the same either side of the border, and two providers score at Ready level on it.

On the country map below, each country takes the weakest rating among its main spoken languages: the most widely spoken one, any home language of roughly a fifth of people, and any shared language routinely used in fieldwork. Of the 22 countries where a main language is still Hard, 20 are in Africa.

World map of countries coloured by the weakest AI transcription rating among their main languages
Each country coloured by the weakest transcription rating among its main spoken languages. Grey: the most widely spoken language isn't rated.

Which provider for which language

“Use AI” isn't one decision. The engines differ a lot, and outside English the difference is not a few percent, it can be several times over. Here is what the independent and semi-independent benchmarks from 2025 and 2026 show when they test the big providers on the same real audio.

Dot chart of speech-to-text error rate by provider across nine languages
Share of words (or characters, for Korean, Japanese and Cantonese) transcribed wrong. Sources: Artificial Analysis (English, Oct 2026), Voice of India (Hindi, Telugu), GigaSpeechBench (others).

What the head-to-heads show

  • On English, it barely matters. On the Artificial Analysis leaderboard, eight big providers' newest models all get between about 2% and 5% of words wrong.
  • On Indian phone calls, it matters a lot. In Voice of India, an independent benchmark of real phone conversations in 15 languages, Sarvam (an Indian specialist) was the best commercial system on 13 of them. Among the big global providers, Google's Gemini 3 Pro led on 11, but Amazon Transcribe won Telugu, Kannada and Odia. OpenAI's GPT-4o Transcribe got 34% of Hindi words wrong against Gemini's 6%, and AssemblyAI's output was unusable in six languages.
  • On Arabic, it depends on the country. In GigaSpeechBench (real YouTube audio, co-written by Alibaba), Google's Chirp 3 was the best Western provider on Saudi Arabic, Google's Gemini on Syrian and Egyptian, and Microsoft Azure on Iraqi. Two Google products swap places from one country to the next.
  • On East and Southeast Asian languages, there's no pattern. Among Western providers, ElevenLabs led on Korean and Thai, Google's Chirp 3 on Vietnamese and Indonesian, and Microsoft on Japanese. On Cantonese, Microsoft got 12% of characters wrong and Google's Chirp 3 got 48%. (Alibaba co-wrote this benchmark, and its own models were best on several of these languages.)
  • On mixed African-English speech, everyone struggles. In AfriSwitch, Gemini and ElevenLabs split the 12 languages six each, averaging 55 to 56% of words wrong. Intron's own model averaged 36%, but Intron also built the benchmark.
  • On African-accented English, Microsoft, Amazon and Google were within about two points of each other on average (22 to 24%), with the winner changing by country. Again an Intron benchmark, where Intron's model was well ahead.
Provider Where the evidence puts it Watch out for
Google (Gemini, Chirp 3) The most consistently strong global provider in the independent tests. Led the big providers on most Indian languages and on several Arabic dialects, and lists more African languages than most. Not uniform: Chirp 3 was weakest of all on Cantonese, and Gemini struggled on Fulani. The two Google engines rank differently by language.
Microsoft (Azure Speech, MAI-Transcribe-2) Best Western provider on Cantonese, Japanese and Iraqi Arabic. MAI-Transcribe-2 is near the top on English. Older Azure was weak on Indian phone audio. Microsoft's claim that MAI-Transcribe-2 is first across 60 languages is its own test.
Amazon Transcribe Close behind Gemini on Indian phone calls, and the best of the big global providers on Telugu, Kannada and Odia. Competitive on African-accented English. Covers fewer languages: no Assamese, Maithili or Urdu in the Indian test, and no isiXhosa, Yoruba or Igbo.
ElevenLabs Scribe v2 Near the top on English, and top on the five big European languages in the Open ASR Leaderboard. Best Western provider on Korean and Thai, and best global provider on Assamese. Publishes accuracy bands for every language. Weak on Saudi Arabic in real audio. Its published bands come from clean audio, so real recordings land worse.
OpenAI (GPT-4o Transcribe, GPT Transcribe) Competitive on English. GPT-4o Transcribe was last or near last on Indian, Arabic and Asian real-world audio. The newer GPT Transcribe hasn't been independently tested outside English yet.
Deepgram, AssemblyAI Strong on English and European languages. Weak or unsupported on most Indian languages; AssemblyAI's Universal model produced unusable output in six of them.
Regional specialists Sarvam (India) was the best commercial system on 13 of 15 Indian languages in an independent test. Intron (Africa) leads its own benchmarks. Lelapa covers several South African languages. Intron's results are self-published, and we found no head-to-head test of Lelapa.
Open models (Meta Omnilingual, Whisper) Cover the long tail nobody sells: Omnilingual handles 1,600+ languages and can be fine-tuned on local audio. You need engineering resources to host and tune them, and zero-shot accuracy on small languages is often poor.
Claude Not a transcription option: the Claude API takes text and images, not audio. Useful once you have a transcript, for translation and analysis.

Two caveats. Model versions change every few months, so any ranking is a snapshot. And the newest models (Gemini 3.5 Transcribe, GPT Transcribe, MAI-Transcribe-2) only have vendor figures outside English so far. The practical answer is to shortlist two or three providers for each language and test them on 20 or 30 of your own voice notes before you commit.

Which model translates best

Translation is closer, but the models still aren't interchangeable. By our calculation from the AI Language Proficiency Monitor's published results, which score frontier models on translation into 194 languages, Gemini 3.1 Pro averaged 52.4, GPT-5.5 50.9 and Claude Opus 4.8 49.6 (the earlier Claude Opus 4.7 scored 51.1). Gemini came top in 124 languages, GPT in 36 and Claude in 35. Each language is scored on only 10 sentences, so gaps of a few points on a single language don't mean much.

Where What the tests found
Sub-Saharan Africa Gemini's clearest lead: 45.9 against 43.5 for GPT and 40.0 for Claude. The gap is wider translating out of these languages.
South African languages The leader changes by language, on small samples: Claude on isiXhosa, Gemini on isiZulu, GPT on siSwati, Google Translate on Xitsonga.
Indian and Southeast Asian languages The three frontier models are statistically tied. Google Translate wins 12 of the 21 Indian languages. It returned nothing for 36 languages, mostly because of a language-code mismatch in the benchmark rather than missing support.
Professional human evaluation (WMT25) Gemini 2.5 Pro was in the top group for 14 of 16 language pairs, and beat the human reference for English to Bhojpuri. But GPT-4.1 clearly beat it on Egyptian Arabic, and Claude 4 beat it on English to Chinese.
Hausa (2026 study) Claude scored best on automatic metrics, but human raters preferred GPT (4.46 out of 5 against 4.19) and put Gemini last (3.37).
43 Ghanaian languages (Nsanku, 2026) Gemini 2.5 Flash 26.9, Claude Sonnet 4.5 24.9, GPT-4.1 23.2. The authors say none is reliably usable yet.
European and global pairs (Alconost, vendor-run, May 2026) Gemini first overall, then Claude, GPT and DeepL. Claude led on German, Brazilian Portuguese and Turkish; DeepL on European Portuguese.

Two things matter more for research than the leaderboard. Benchmarks use formal text, and in WMT25 transcribed speech was the hardest material to translate, likely because of errors carried over from the transcription. And automatic scores and human judges don't always agree, as the Hausa study shows. For open-ended answers, have a native speaker read a sample of translations alongside the originals, especially where people use slang or switch languages.

Accents inside “ready” languages

“Ready” is a rating for a whole language, and it hides a lot. On standard test sets the best systems get about 2% of English words wrong. On real people, the picture changes depending on where they're from and who they are.

Dumbbell chart comparing speech-to-text error rate on standard English and on regional or community speech
Each line compares one system on two groups of speakers in the same study. Sources: VarDial 2025; Markl, FAccT 2022; Koenecke et al., PNAS 2020; Harris et al., EMNLP Findings 2024.
  • Scotland. On real Scottish speech in GigaSpeechBench, the best system tested got 24% of words wrong (an Alibaba model, in a benchmark Alibaba co-wrote). Microsoft Azure got 28%, and ElevenLabs, Gemini and GPT-4o were between 34% and 44%. Whisper went from 4% on its baseline to 22 to 34% on Scottish housing and charity calls.
  • Northern England and Northern Ireland. On Newcastle speech, the best of four systems averaged 32%, with Google and Deepgram at 40 to 60%. In an older study of nine British cities, Amazon Transcribe was 9 points worse in Belfast and Bradford than in Cambridge.
  • Black American speakers. In the reference study from 2020, five commercial systems averaged 35% against 19% for white speakers, and 23% of clips were unusable against 1.6%. The systems are older, and we found no repeat of the test on today's models.
  • Spanglish and Chicano English. An open model got 42% and 38% of words wrong, against 20.5% on standard American English. We found no independent test of today's commercial providers on real Spanglish conversation.
  • Varieties of the same language. A model tuned on European Portuguese made 27% errors on Brazilian Portuguese against 12.5% on its own variety. Southern Spain accents ran up to 8 points worse than northern ones. Québec French came out at 8 to 10%, but standard French benchmarks didn't predict which model would do best.

The home languages of UK and US samples matter too. Only 78.4% of Londoners have English as their main language (91.1% across England and Wales). The next biggest languages nationally are Polish, Romanian, Panjabi and Urdu. Polish and Romanian are Ready. Panjabi and Urdu are Ready only on evidence from Indian audio (Sarvam and Gemini). ElevenLabs rates Urdu at 25 to 50% of words wrong, Google's Chirp 3 has no Pakistani Urdu locale, and we found no test on UK speakers. In the US, Spanish is spoken at home by 13.9% of people, followed by Chinese, Tagalog, Vietnamese and Arabic, all Ready, while Haitian Creole, with about a million speakers, is only Workable.

What to do. Set the regional model where one exists, not just the language (es-MX or es-ES, pt-BR or pt-PT), and test it on your own audio. And spot-check transcripts by region or community, not just by language. A sample that's 95% “English” can still have a group whose transcripts are a third wrong.

Markets we haven't mapped

Our 23-market map leaves out some markets that come up in almost every multi-country brief. Here is where their main languages stand.

Market Main language Where it stands
France, Italy, Poland French, Italian, Polish Ready. ElevenLabs rates all three Excellent and Google has full support.
Vietnam Vietnamese Ready. Best real-world error about 10% (Google's Chirp 3).
South Korea Korean Ready. Best real-world error about 10% of characters (an Alibaba model); ElevenLabs 12%.
Japan Japanese Ready on clean audio, but real-world error was 28 to 44% of characters across Western providers. Part of that is scoring, since the same word can be written correctly in different scripts. Check a sample.
China Mandarin Ready. Regional varieties such as Wu and Xiang are much weaker, and mainly handled by Chinese providers.
Indonesia Indonesian Ready, but about 74% of people also use a regional language, and Javanese is only Workable and Sundanese Hard.
Saudi Arabia, UAE Gulf Arabic Workable. Best real-world error 17% on Saudi audio (Alibaba and Google's Chirp 3 within a point) and 26% on Emirati audio.
Egypt Egyptian Arabic Workable. Best real-world error 37% (an Alibaba model); the best Western provider, Gemini, 41%.
Australia English Ready. We found no accent-specific evidence, and more than 5.5 million people use another language at home.

Code-switching

Real people don't speak one clean language at a time. A respondent in Nairobi slides between Swahili, English and Sheng in a single thought. In Johannesburg it's English threaded through isiZulu. In Lagos, English, Yoruba and Pidgin. In Manila, Taglish. In India, Hinglish.

Most speech models assume one language per sentence, and the gap shows. AfriSwitch, a 2026 benchmark of 61 hours of real conversation mixing English with 16 African languages and varieties, found no system below 24% on any of the 12 languages benchmarked so far. The best, Intron's own Sahara V2.5, averaged 36% across the 12 languages in its main results; Meta's Omnilingual, Google's Gemini and ElevenLabs averaged 51 to 56%. In the Philippines, where Taglish is everyday speech, the best real-world results for Tagalog were still around a quarter of words wrong.

Two things help. Use a speech model trained on mixed speech where one exists (Sarvam, for example, supports Hinglish). And don't assume an LLM can tidy it up afterwards: asking GPT-4o-mini to correct Hindi transcripts without examples took word error from 18% to 25%. Keep the original audio, and have a native speaker check the parts where people switch.

Recording conditions

Benchmark accuracy is measured on clean recordings. Your data is a voice note recorded on a cheap phone in a taxi or a market. Expect error rates roughly 1.5 to 2 times worse than the clean figures for single-language speech. On real Indian phone calls, the noisiest quarter of recordings scored about 1.65 to 1.75 times worse than the cleanest; on African conversational audio the gap ranged from 1.3 to 2.8 times.

The cause matters. WhatsApp's own audio compression (Opus) barely moves accuracy. Old-style 8 kHz phone lines do. And competing voices in the background do the most damage of all. So the cheapest improvement isn't a better model, it's asking people to record somewhere quiet, away from the TV and other conversations, and testing models on your own pilot audio rather than their published numbers.

So, AI or a team?

Put it together and it isn't “AI” or “humans”. It's a split that depends on the language, the provider and the people.

Rating How to set it up
Ready transcription, Strong translation Run it with AI. Spot-check a sample, by region or community if your sample is mixed.
Ready or Workable transcription, Strong or Moderate translation AI drafts, a native speaker checks the discussion guide before launch and the transcripts before analysis.
Hard transcription, but Strong or Moderate translation Ask for typed answers, which AI can read and translate well, or have a person transcribe.
Basic translation Keep a person in the loop for both transcription and translation. AI makes them faster, it doesn't replace them.
01

Check the country, not just the language

Look at the home languages of the people you'll actually reach, and how many speak a Ready one.

02

Pick the engine per language

Shortlist two or three providers for each language and test them on your own pilot audio.

03

Set the variety

Choose the regional model where one exists, and budget a check by region or community.

04

Plan for mixing and noise

Ask for quiet recordings, keep the original audio, and have native speakers check the switched parts.

The right answer is AI-first, with people placed exactly where the technology is weak.

That's how research happens in people's own words at scale without quietly shipping bad data. “We can only do this in English” is no longer true for most large languages. Which languages, in which mode, with how much human review, is now a decision worth making on purpose.

Frequently asked questions

Can AI accurately transcribe African languages?

Some of them. Swahili and Afrikaans are the only two of the 44 sub-Saharan languages we rated that come out Ready. Most of the big ones, like isiZulu, Hausa, Yoruba and Amharic, are Workable: AI gives you a usable draft and a native speaker needs to check it. A long tail, including Tshivenda, isiNdebele and many smaller languages, still has no modern commercial speech-to-text. Mixing languages makes all of them harder.

Which AI transcription provider is best?

None of them across the board. On English the big providers all sit between about 2% and 5% of words wrong. Outside English the winner changes by language: on Indian phone calls Google's Gemini led on most languages but Amazon won Telugu, Kannada and Odia; on Cantonese Microsoft made about a quarter as many errors as Google's Chirp 3. Regional specialists such as Sarvam in India often beat all of them. Test two or three on your own audio before you commit.

Can Claude or ChatGPT transcribe interviews?

Claude can't. Its API takes text and images, not audio, so it isn't a transcription option, although it works well for translating and analysing a transcript once you have one. OpenAI does sell transcription models. Its GPT-4o Transcribe model was competitive on English but among the weakest on Indian phone audio and real-world Arabic and Asian audio in the independent tests we found. Its newer model hasn't been tested independently outside English yet.

Is English transcription solved?

On standard test sets, nearly. On real people, not quite. In the studies we found, error rates roughly doubled for Northern Irish and Northern English speakers, Black American speakers and Spanglish, and went up several times over for Scottish speech. If your sample spans regions or communities, check a few transcripts from each group, not just a few in total.

Why is translation usually easier than transcription?

Translation starts from text, and large language models have read far more text than speech models have heard audio. For about one language in five, the translation rating is a tier ahead of the transcription rating, and it's rarely behind. The catch is informal speech: in the 2025 machine translation shared task, transcribed speech was the hardest material to translate.

Should I let an LLM clean up a messy transcript?

Carefully. In a 2026 study, asking GPT-4o-mini to correct Hindi transcripts without examples pushed the word error rate from 18% to 25%. With good examples it helped in some setups and not others. Use a speech model trained on mixed speech where one exists, keep the original audio, and have a native speaker check the parts where people switch languages.

Do researchers still need human transcribers in 2026?

Yes, but in a smaller and more targeted role. For Ready languages a light spot-check is enough. For Workable languages a native speaker should check the discussion guide before launch and the transcripts before analysis. For Hard languages, ask for typed answers or have a person lead the transcription with AI helping.

Running something multilingual?

That's our home ground.

We'll tell you which languages are ready to run on AI and which ones need a human hand, market by market.

Book a Demo →

Related Posts