You have research to run in Swahili, or isiZulu, or Hindi, or across a dozen markets at once, and you want one answer: can AI transcribe and translate it, or do you still need people doing it by hand? This is our language-by-language answer, rebuilt in October 2026 with 130 languages, 23 markets and a head-to-head of the main providers.
How the ratings work
The language table
Languages inside each country
Which provider for which language
Which model translates best
Accents inside “ready” languages
Markets we haven't mapped
Code-switching
Recording conditions
So, AI or a team?
FAQ
The short version: for English, Spanish and the big European and Asian languages, AI now does the work and a light check is enough. For most of Africa's large languages, AI gives you a good draft that a native speaker needs to check. For a long tail, it still can't hear the language at all. And two things cut across every rating: which provider you use, and who is speaking. The same set of Hindi phone calls can come back with 6% of words wrong or 34%, depending on the engine. The same system can get English near perfect when it's read into a studio microphone, and a third wrong on calls to an advice charity in North East Scotland.
What changed since August
We published the first version of this guide in July and updated it in August. When we checked it again in October, 23 of the 74 ratings had moved, and we'd missed a lot of languages people actually field in. So we rebuilt it against what the providers support today and the newest real-world benchmarks.
- Up. Cantonese, Punjabi, Odia, Assamese and Nepali are now Ready, and so is Levantine Arabic, though that rests on one benchmark. Bhojpuri, Luganda, Wolof and Nigerian Pidgin moved from Hard to Workable. Pidgin finally has a commercial speech model (Intron Sahara v2, March 2026). Translation into isiZulu, Kinyarwanda and Javanese is now Strong, and Xitsonga, siSwati and Kirundi moved up to Moderate.
- Down. Hausa and Egyptian Arabic dropped to Workable once we looked at real recordings rather than clean test audio. Moroccan Darija, Sudanese Arabic and Kurmanji Kurdish dropped to Hard. Hebrew is now a borderline Workable. Translation into Wolof and Fula dropped to Basic.
- New. 56 languages added, including Turkish, Malay, Gulf and Iraqi Arabic, Pashto, Dari, Central Asian languages, 14 more sub-Saharan African languages, Tamazight and Kabyle, Haitian Creole, and indigenous languages of Latin America such as Quechua, Aymara, Guarani, Nahuatl and the main Mayan languages.
- Advice we've changed. We used to suggest transcribing first and letting an LLM clean up mixed-language text. A 2026 study found that made Hindi transcripts worse. We also used to point to Meta's MMS model for the long tail. Meta has since replaced it with Omnilingual ASR (November 2025, open source, 1,600+ languages), though it still leaves out Sesotho, siSwati, Tshivenda and isiNdebele.
- New sections. Which provider wins where, which model translates best, how accents change the picture inside “ready” languages, and the language mix inside 23 markets.
How the ratings work
Two separate ratings, because they're two separate jobs. Transcription turns a voice note into text in the same language. Translation turns that text into another language. Each language is rated on the best tool available on the market, not on any one provider, and we apply the same rules to every language.
-
●
Ready (transcription). At least two commercial providers, and either ElevenLabs rates its own accuracy at 10% word error or better, or the best real-world benchmark is around 15% or better.
-
◐
Workable. At least one current commercial provider, including regional specialists such as Intron, Sarvam and Lelapa, with real-world error of roughly 15 to 35%, or a listing with no published accuracy.
-
○
Hard. No current commercial speech-to-text, or best evidence worse than about 35 to 40% of words wrong.
-
●
Strong / Moderate / Basic (translation). Based on how closely frontier models (GPT, Claude and Gemini) match a reference translation into the language, on the AI Language Proficiency Monitor's 0 to 100 chrF score: Strong at 57 or above, Moderate from about 42 to 56 or listed by Google Translate or DeepL with no benchmark, Basic below that. We rate up where the score is known to undercount, such as Ge'ez script (Amharic, Tigrinya) and Yoruba tone marks.
The language table
All 130 languages, grouped by region. The last column tells you what to do in fieldwork and the evidence behind the rating. Moved marks a rating that changed since August, New a language we added. Speaker counts are broad public estimates.
● Ready AI does the job, light spot-check. ◐ Workable AI drafts, a native speaker checks. ○ Hard Human-led, AI assists.
| Language | Speakers | Transcription | Translation | What to do, and the evidence |
|---|---|---|---|---|
| Southern Africa | ||||
| Afrikaans | 12M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Good (10-20%); open-model bench 9.6%; chrF 68 |
| isiZuluMoved | 28M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Translation up: chrF 61, top-ranked LLMs in WMT24++. ASR 16-26% (AfriVox), 48%+ code-switched |
| isiXhosa | 19M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; AfriVox 17-29%; chrF 52 |
| Sesotho | 14M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Commercial via Google legacy, Intron, Lelapa; AfriVox best 15%; not in Meta Omnilingual |
| Sepedi | 14M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; Chirp 3 preview (nso-ZA); chrF 53 (thin evidence) |
| Setswana | 14M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google legacy + Intron; AfriVox best 11.6% (vendor bench); chrF 48 |
| Shona | 15M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; AfriVox 21-37% |
| Chichewa | 20M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google chirp_2 only; best 22%, most results 30%+ (thin evidence) |
| XitsongaMoved | 3.5M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Translation up: chrF 56, Google Translate 63. ASR: Google legacy model only |
| siSwatiMoved | 3M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Translation up: Google Translate added 2024, chrF 51. ASR: no modern commercial model |
| Tshivenda | 2M | ○ Hard | ○ Basic | Human-led, AI assists.Google legacy ASR only; translation consumer-app only, no benchmark (thin evidence) |
| isiNdebele | 1.5M | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; open models only (thin evidence) |
| East Africa | ||||
| Swahili | 95M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High (5-10%); AfriVox 7%; chrF 70 |
| Amharic | 78M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AfriVox ~25%; chrF 37 into / 55 out of Amharic (script undercounts) |
| Somali | 22M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AWS, Azure, Google; ElevenLabs Moderate; open bench 35% (thin evidence) |
| KinyarwandaMoved | 15M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Translation up: chrF 60. ASR mixed: 6.6% (Intron, own bench) to 30%+ elsewhere |
| LugandaMoved | 10M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: AWS, Google, ElevenLabs, Intron now commercial; best 16% (Sunbird) |
| Oromo | 46M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google chirp_2, Intron; open bench 21-25%; chrF 43 |
| Tigrinya | 9M | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; chrF 30 into Tigrinya |
| KirundiMoved | 12M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Translation up: Google Translate added 2024, chrF 49. No commercial ASR |
| DholuoNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google (luo-KE) and ElevenLabs list it; open bench 17.5% (thin evidence) | |
| KikuyuNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR (open bench 22.6%); not in Google Translate | |
| MalagasyNew | ◐ Workable | ● Strong | AI drafts, a native speaker checks.AssemblyAI and Gladia list it; open bench 13.7%; chrF 57 (thin evidence) | |
| DinkaNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; Google Translate 2024 (thin evidence) | |
| West & Central Africa | ||||
| HausaMoved | 94M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Transcription down: ElevenLabs Good (10-20%); best real-world 17-27%; 32% code-switched |
| Yoruba | 50M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AfriVox 27-36%; 45%+ code-switched; translation drops tone marks (chrF 33) |
| Igbo | 30M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; best 24%; chrF 53 |
| Lingala | 40M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google chirp_2, AssemblyAI; open bench 21.5% (thin evidence) |
| Akan (Twi) | 20M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Intron; best 28.7%; chrF 46 |
| WolofMoved | 12M | ◐ Workable | ○ Basic | Human-led, AI assists.Transcription up (AWS, Chirp 3 preview, Intron; best ~28-32%). Translation down: chrF 34 |
| FulaMoved | 40M | ○ Hard | ○ Basic | Human-led, AI assists.Translation down: chrF 25. ASR best 34%, most results 40-70% |
| Nigerian PidginMoved | 60-120M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: Intron Sahara v2 supports it (Mar 2026); open bench 19.7%. Not in Google/Azure (thin evidence) |
| Mossi | 9M | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; not in Google Translate; chrF 29 |
| Kanuri | 9M | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; open bench 35%; Google Translate app only (thin evidence) |
| EweNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No commercial ASR (open bench 29%); Google Translate (thin evidence) | |
| BambaraNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; open bench 44%; chrF 34 | |
| FonNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No commercial ASR; Google Translate 2024 (thin evidence) | |
| KrioNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No commercial ASR; English-based creole, Google Translate (thin evidence) | |
| SangoNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No commercial ASR; Google Translate 2024 (thin evidence) | |
| BembaNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Intron lists it (no figure); open bench 37-44%; chrF 48 | |
| TshilubaNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; NLLB only (thin evidence) | |
| MakhuwaNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; no Google Translate (thin evidence) | |
| UmbunduNew | ◐ Workable | ○ Basic | Human-led, AI assists.Listed by ElevenLabs and Google with no accuracy figure; not in Google Translate (thin evidence) | |
| OshiwamboNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; no Google Translate (thin evidence) | |
| Middle East & North Africa | ||||
| Arabic (MSA) | 335M | ● Ready | ● Strong | AI, light spot-check.Broad coverage; ElevenLabs Good; best open model 10.2% on MSA |
| Egyptian ArabicMoved | 118M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Transcription down: best real-world 37% (Alibaba, GigaSpeechBench); best Western provider Gemini 41%, ElevenLabs 44% |
| Levantine ArabicMoved | 58M | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: real-world Syria best 13.8% (Alibaba), Gemini 14.4%, ElevenLabs 14.7% (thin evidence) |
| Gulf ArabicNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Google, Azure, AWS (ar-AE), Deepgram; real-world Saudi best 16.6% (Alibaba), Google Chirp 3 16.8%; UAE 26% | |
| Iraqi ArabicNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Azure, Deepgram, Google ar-IQ; real-world best 28.5% (Alibaba), Microsoft Azure 34.6% | |
| Moroccan DarijaMoved | 21M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Transcription down: best real-world 51%, every system above 44% |
| Algerian & Tunisian ArabicNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Locales exist (Azure, Google) but best real-world 44% (thin evidence) | |
| Sudanese ArabicMoved | 54M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Transcription down: Deepgram ar-SD only; best open model 36% (thin evidence) |
| TamazightNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; Google Translate 2024 (Latin and Tifinagh) (thin evidence) | |
| KabyleNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AWS Transcribe (kab-DZ), no accuracy published; open 29% (thin evidence) | |
| Persian | 82M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High (5-10%); broad coverage |
| Kurdish (Kurmanji)Moved | 26M | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.Transcription down: Kurmanji not supported by any major provider (Sorani only) |
| Kurdish (Sorani)New | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AWS and Google (ckb-IQ, ckb-IR); ElevenLabs 'Kurdish' Moderate (thin evidence) | |
| HebrewMoved | 9M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Transcription down to the rule: ElevenLabs Good (10-20%), no real-world benchmark. Borderline (thin evidence) |
| TurkishNew | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent (5% or under); every major provider | |
| South Asia | ||||
| Hindi | 611M | ● Ready | ● Strong | AI, light spot-check.Real-world 5-8% (Voice of India) |
| Bengali | 274M | ● Ready | ● Strong | AI, light spot-check.Real-world 6-10% |
| Urdu | 246M | ● Ready | ● Strong | AI, light spot-check.Indian phone audio: Sarvam 7.0%, Gemini 3 Pro 9.1%. ElevenLabs rates it Moderate (25 to 50%); Google Chirp 3 has no Pakistani Urdu locale |
| Tamil | 86M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High; real-world best 14% |
| Telugu | 96M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High; real-world best 18% |
| Marathi | 99M | ● Ready | ● Strong | AI, light spot-check.Real-world 9-11% |
| Gujarati | 62M | ● Ready | ● Strong | AI, light spot-check.Real-world best 12.8% |
| Kannada | 59M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent; real-world best 16% |
| Malayalam | 38M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent; real-world best 19% |
| PunjabiMoved | 90M | ● Ready | ● Strong | AI, light spot-check.Transcription up: Indian phone audio best 11.2% (Sarvam), Gemini 14.4%. Pakistani Punjabi (Shahmukhi) not benchmarked; treat as Workable in Pakistan |
| OdiaMoved | 38M | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: ElevenLabs High; real-world best 14% |
| AssameseMoved | 24M | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: real-world best 12.7% |
| NepaliMoved | 32M | ● Ready | ● Strong | AI, light spot-check.Transcription up: ElevenLabs High (5-10%) |
| Sindhi | 37M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; Sarvam covers it |
| Sinhala | 26M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; Azure, AWS |
| BhojpuriMoved | 53M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Transcription up: real-world best 18.4% (Gemini 3 Pro), Sarvam 20.9% |
| MaithiliNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Sarvam covers it; real-world best 24.7% | |
| PashtoNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate (25-50%); AWS, Azure, Deepgram, AssemblyAI | |
| DariNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.AWS fa-AF; otherwise relies on Persian models (thin evidence) | |
| SaraikiNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; not in Google Translate (thin evidence) | |
| East & Southeast Asia | ||||
| Mandarin | 1,184M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High |
| Japanese | 126M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent on clean audio, but real-world 28 to 44% of characters wrong across Western providers (partly scoring: kanji vs kana) |
| Korean | 82M | ● Ready | ● Strong | AI, light spot-check.Real-world best 9.9% of characters wrong (Alibaba FunASR, GigaSpeechBench); ElevenLabs 11.8%, Microsoft 13.1% |
| Indonesian | 255M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent; real-world best 15% (Alibaba), Google Chirp 3 20% |
| Vietnamese | 97M | ● Ready | ● Strong | AI, light spot-check.Real-world best 9.6% (Google Chirp 3) |
| Thai | 71M | ● Ready | ● Strong | AI, light spot-check.Real-world best 10.8% (Alibaba), ElevenLabs 13.9% |
| Tagalog | 87M | ● Ready | ● Strong | AI, light spot-check.ElevenLabs High; with Taglish mixing, real-world best about a quarter of words wrong |
| MalayNew | ● Ready | ● Strong | AI, light spot-check.ElevenLabs Excellent; every major provider | |
| CantoneseMoved | 86M | ● Ready | ● Strong | AI, light spot-check.Real-world best 6.1% CER (Alibaba FunASR); Microsoft Azure 11.8%, Google Chirp 3 48% |
| Cebuano | 28M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Listed by ElevenLabs, Google, Gemini with no accuracy figure (thin evidence) |
| JavaneseMoved | 69M | ◐ Workable | ● Strong | AI drafts, a native speaker checks.Translation up: chrF 60. ASR ElevenLabs Good |
| Burmese | 43M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Good; chrF 56 |
| Khmer | 18M | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate |
| LaoNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; Google, Azure | |
| SundaneseNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No major commercial ASR; Google, DeepL translate it (thin evidence) | |
| Taiwanese HokkienNew | ○ Hard | ○ Basic | Human-led, AI assists.ElevenLabs 71%, Azure 67% real-world; specialist open models ~28-30% CER | |
| Central Asia & Caucasus | ||||
| AzerbaijaniNew | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure (thin evidence) | |
| KazakhNew | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram. Russian mixing common | |
| UzbekNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Good (10-20%); Google, AWS, Azure | |
| KyrgyzNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Good; Google preview (thin evidence) | |
| TajikNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Good; not Azure or Chirp 3 (thin evidence) | |
| GeorgianNew | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram | |
| ArmenianNew | ● Ready | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs High; Google, AWS, Azure, Deepgram | |
| MongolianNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate (25-50%); Google, AWS, Azure | |
| Americas & Caribbean | ||||
| Haitian CreoleNew | ◐ Workable | ● Strong | AI drafts, a native speaker checks.AWS Transcribe (ht-HT) and AssemblyAI list it, no accuracy published; chrF 57 (thin evidence) | |
| QuechuaNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; translation weak (AmericasNLP) | |
| GuaraniNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR; chrF 39 | |
| AymaraNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial ASR | |
| NahuatlNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text; Google Translate app only (Eastern Huasteca) | |
| Yucatec MayaNew | ○ Hard | ◐ Moderate | Ask for typed answers, or a human transcribes.No commercial speech-to-text; Google Translate incl. Cloud API (thin evidence) | |
| TseltalNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| TsotsilNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| MixtecNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| ZapotecNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text; Google Translate app only | |
| K'iche'New | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| Q'eqchi'New | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text; Google Translate app only | |
| MamNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text; Google Translate app only | |
| KaqchikelNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| KichwaNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text; translation only via generic Quechua | |
| MapudungunNew | ○ Hard | ○ Basic | Human-led, AI assists.No commercial speech-to-text or machine translation | |
| Europe & global | ||||
| English | 1.5B | ● Ready | ● Strong | AI, light spot-check.Leaderboard 2 to 5% across all big providers; regional accents far higher (real Scottish speech, best system 24%) |
| Spanish | 561M | ● Ready | ● Strong | AI, light spot-check.Choose the regional model (es-MX, es-US, es-ES). Spanglish roughly doubled errors in a 2024 study of open models |
| French | 334M | ● Ready | ● Strong | AI, light spot-check.Québec French 8 to 10% in a 2025 test; standard French benchmarks didn't predict which model did best |
| Portuguese | 269M | ● Ready | ● Strong | AI, light spot-check.Choose pt-BR or pt-PT: a model tuned on European Portuguese made twice the errors on Brazilian |
| Russian | 210M | ● Ready | ● Strong | AI, light spot-check. |
| German | 133M | ● Ready | ● Strong | AI, light spot-check. |
| Italian | 66M | ● Ready | ● Strong | AI, light spot-check. |
| Other European | ● Ready | ● Strong | AI, light spot-check.Serbian checked: ElevenLabs High | |
| AlbanianNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.Azure only; not in ElevenLabs bands or Chirp 3 (thin evidence) | |
| WelshNew | ◐ Workable | ● Strong | AI drafts, a native speaker checks.ElevenLabs Good; Azure, Speechmatics | |
| IrishNew | ◐ Workable | ◐ Moderate | AI drafts, a native speaker checks.ElevenLabs Moderate; Azure, Speechmatics | |
About 36 of these ratings rest on thin evidence, mostly newer African languages where all we have is a provider listing or an open-model test. They're marked in the table. The ratings move quickly, so check the date on anything you rely on.
Languages inside each country
A language rating only goes so far, because nobody runs a study in “isiZulu”. They run it in South Africa, where English is spoken at home by fewer than one in ten people. So we mapped the home languages of 23 markets, using censuses and statistics-office surveys, and coloured each language by its transcription rating.
The split is stark. In Brazil, the UK, the US, Mexico and Spain, more than nine in ten people speak a home language AI transcribes well. In Ethiopia, Uganda, Nigeria, Ghana and Pakistan, fewer than one in ten do. South Africa sits at about one in five, because English and Afrikaans are Ready but isiZulu, isiXhosa, Sepedi, Setswana and Sesotho are only Workable.
| Market | Home language rated Ready | Largest home languages that are not Ready |
|---|---|---|
| Brazil | >99% | None above 1% of people |
| United Kingdom | 96% | None above 1% of people |
| United States | 95% | None above 1% of people |
| Mexico | 94% | Tseltal, Tsotsil, Mixtec, Zapotec 2%, Nahuatl 1% |
| Spain | 93% | None above 1% of people |
| Canada | 90% | None above 1% of people |
| Germany | 88% | None above 1% of people |
| Turkey | 86% | Kurdish/Zazaki 13% |
| Peru | 83% | Quechua 14%, Aymara 2% |
| Bolivia | 78% | Quechua 13%, Aymara 7% |
| India | 76% | Bhojpuri 4%, Maithili 1% |
| Tanzania | 76% | Sukuma 8% (not rated) |
| Guatemala | 70% | Q'eqchi' 8%, K'iche' 8%, Mam 4% |
| Nepal | 46% | Maithili 11%, Bhojpuri 6%, Tharu 6% (not rated) |
| Philippines | 40% | Bisaya/Cebuano 22%, Hiligaynon 7% (not rated), Ilocano 7% (not rated) |
| Kenya | 33% | Kikuyu 13%, Kalenjin 8% (not rated), Kamba 8% (not rated) |
| Paraguay | 31% | Guarani + Spanish 39% (both), Guarani 28% |
| South Africa | 19% | isiZulu 24%, isiXhosa 16%, Sepedi 10% |
| Pakistan | 9% | Punjabi 37%, Pashto 18%, Sindhi 14% |
| Ghana | 7% | Akan 52%, Ewe 12%, Dagbani 7% (not rated) |
| Nigeria | 6% | Hausa 32%, Yoruba 17%, Igbo 11% |
| Uganda | 3% | Luganda 28%, Runyankore 10% (not rated), Lusoga 10% (not rated) |
| Ethiopia | 0% | Oromo 34%, Amharic 29%, Somali 6% |
On the country map below, each country takes the weakest rating among its main spoken languages: the most widely spoken one, any home language of roughly a fifth of people, and any shared language routinely used in fieldwork. Of the 22 countries where a main language is still Hard, 20 are in Africa.
Which provider for which language
“Use AI” isn't one decision. The engines differ a lot, and outside English the difference is not a few percent, it can be several times over. Here is what the independent and semi-independent benchmarks from 2025 and 2026 show when they test the big providers on the same real audio.
What the head-to-heads show
- On English, it barely matters. On the Artificial Analysis leaderboard, eight big providers' newest models all get between about 2% and 5% of words wrong.
- On Indian phone calls, it matters a lot. In Voice of India, an independent benchmark of real phone conversations in 15 languages, Sarvam (an Indian specialist) was the best commercial system on 13 of them. Among the big global providers, Google's Gemini 3 Pro led on 11, but Amazon Transcribe won Telugu, Kannada and Odia. OpenAI's GPT-4o Transcribe got 34% of Hindi words wrong against Gemini's 6%, and AssemblyAI's output was unusable in six languages.
- On Arabic, it depends on the country. In GigaSpeechBench (real YouTube audio, co-written by Alibaba), Google's Chirp 3 was the best Western provider on Saudi Arabic, Google's Gemini on Syrian and Egyptian, and Microsoft Azure on Iraqi. Two Google products swap places from one country to the next.
- On East and Southeast Asian languages, there's no pattern. Among Western providers, ElevenLabs led on Korean and Thai, Google's Chirp 3 on Vietnamese and Indonesian, and Microsoft on Japanese. On Cantonese, Microsoft got 12% of characters wrong and Google's Chirp 3 got 48%. (Alibaba co-wrote this benchmark, and its own models were best on several of these languages.)
- On mixed African-English speech, everyone struggles. In AfriSwitch, Gemini and ElevenLabs split the 12 languages six each, averaging 55 to 56% of words wrong. Intron's own model averaged 36%, but Intron also built the benchmark.
- On African-accented English, Microsoft, Amazon and Google were within about two points of each other on average (22 to 24%), with the winner changing by country. Again an Intron benchmark, where Intron's model was well ahead.
| Provider | Where the evidence puts it | Watch out for |
|---|---|---|
| Google (Gemini, Chirp 3) | The most consistently strong global provider in the independent tests. Led the big providers on most Indian languages and on several Arabic dialects, and lists more African languages than most. | Not uniform: Chirp 3 was weakest of all on Cantonese, and Gemini struggled on Fulani. The two Google engines rank differently by language. |
| Microsoft (Azure Speech, MAI-Transcribe-2) | Best Western provider on Cantonese, Japanese and Iraqi Arabic. MAI-Transcribe-2 is near the top on English. | Older Azure was weak on Indian phone audio. Microsoft's claim that MAI-Transcribe-2 is first across 60 languages is its own test. |
| Amazon Transcribe | Close behind Gemini on Indian phone calls, and the best of the big global providers on Telugu, Kannada and Odia. Competitive on African-accented English. | Covers fewer languages: no Assamese, Maithili or Urdu in the Indian test, and no isiXhosa, Yoruba or Igbo. |
| ElevenLabs Scribe v2 | Near the top on English, and top on the five big European languages in the Open ASR Leaderboard. Best Western provider on Korean and Thai, and best global provider on Assamese. Publishes accuracy bands for every language. | Weak on Saudi Arabic in real audio. Its published bands come from clean audio, so real recordings land worse. |
| OpenAI (GPT-4o Transcribe, GPT Transcribe) | Competitive on English. | GPT-4o Transcribe was last or near last on Indian, Arabic and Asian real-world audio. The newer GPT Transcribe hasn't been independently tested outside English yet. |
| Deepgram, AssemblyAI | Strong on English and European languages. | Weak or unsupported on most Indian languages; AssemblyAI's Universal model produced unusable output in six of them. |
| Regional specialists | Sarvam (India) was the best commercial system on 13 of 15 Indian languages in an independent test. Intron (Africa) leads its own benchmarks. Lelapa covers several South African languages. | Intron's results are self-published, and we found no head-to-head test of Lelapa. |
| Open models (Meta Omnilingual, Whisper) | Cover the long tail nobody sells: Omnilingual handles 1,600+ languages and can be fine-tuned on local audio. | You need engineering resources to host and tune them, and zero-shot accuracy on small languages is often poor. |
| Claude | Not a transcription option: the Claude API takes text and images, not audio. | Useful once you have a transcript, for translation and analysis. |
Two caveats. Model versions change every few months, so any ranking is a snapshot. And the newest models (Gemini 3.5 Transcribe, GPT Transcribe, MAI-Transcribe-2) only have vendor figures outside English so far. The practical answer is to shortlist two or three providers for each language and test them on 20 or 30 of your own voice notes before you commit.
Which model translates best
Translation is closer, but the models still aren't interchangeable. By our calculation from the AI Language Proficiency Monitor's published results, which score frontier models on translation into 194 languages, Gemini 3.1 Pro averaged 52.4, GPT-5.5 50.9 and Claude Opus 4.8 49.6 (the earlier Claude Opus 4.7 scored 51.1). Gemini came top in 124 languages, GPT in 36 and Claude in 35. Each language is scored on only 10 sentences, so gaps of a few points on a single language don't mean much.
| Where | What the tests found |
|---|---|
| Sub-Saharan Africa | Gemini's clearest lead: 45.9 against 43.5 for GPT and 40.0 for Claude. The gap is wider translating out of these languages. |
| South African languages | The leader changes by language, on small samples: Claude on isiXhosa, Gemini on isiZulu, GPT on siSwati, Google Translate on Xitsonga. |
| Indian and Southeast Asian languages | The three frontier models are statistically tied. Google Translate wins 12 of the 21 Indian languages. It returned nothing for 36 languages, mostly because of a language-code mismatch in the benchmark rather than missing support. |
| Professional human evaluation (WMT25) | Gemini 2.5 Pro was in the top group for 14 of 16 language pairs, and beat the human reference for English to Bhojpuri. But GPT-4.1 clearly beat it on Egyptian Arabic, and Claude 4 beat it on English to Chinese. |
| Hausa (2026 study) | Claude scored best on automatic metrics, but human raters preferred GPT (4.46 out of 5 against 4.19) and put Gemini last (3.37). |
| 43 Ghanaian languages (Nsanku, 2026) | Gemini 2.5 Flash 26.9, Claude Sonnet 4.5 24.9, GPT-4.1 23.2. The authors say none is reliably usable yet. |
| European and global pairs (Alconost, vendor-run, May 2026) | Gemini first overall, then Claude, GPT and DeepL. Claude led on German, Brazilian Portuguese and Turkish; DeepL on European Portuguese. |
Two things matter more for research than the leaderboard. Benchmarks use formal text, and in WMT25 transcribed speech was the hardest material to translate, likely because of errors carried over from the transcription. And automatic scores and human judges don't always agree, as the Hausa study shows. For open-ended answers, have a native speaker read a sample of translations alongside the originals, especially where people use slang or switch languages.
Accents inside “ready” languages
“Ready” is a rating for a whole language, and it hides a lot. On standard test sets the best systems get about 2% of English words wrong. On real people, the picture changes depending on where they're from and who they are.
- Scotland. On real Scottish speech in GigaSpeechBench, the best system tested got 24% of words wrong (an Alibaba model, in a benchmark Alibaba co-wrote). Microsoft Azure got 28%, and ElevenLabs, Gemini and GPT-4o were between 34% and 44%. Whisper went from 4% on its baseline to 22 to 34% on Scottish housing and charity calls.
- Northern England and Northern Ireland. On Newcastle speech, the best of four systems averaged 32%, with Google and Deepgram at 40 to 60%. In an older study of nine British cities, Amazon Transcribe was 9 points worse in Belfast and Bradford than in Cambridge.
- Black American speakers. In the reference study from 2020, five commercial systems averaged 35% against 19% for white speakers, and 23% of clips were unusable against 1.6%. The systems are older, and we found no repeat of the test on today's models.
- Spanglish and Chicano English. An open model got 42% and 38% of words wrong, against 20.5% on standard American English. We found no independent test of today's commercial providers on real Spanglish conversation.
- Varieties of the same language. A model tuned on European Portuguese made 27% errors on Brazilian Portuguese against 12.5% on its own variety. Southern Spain accents ran up to 8 points worse than northern ones. Québec French came out at 8 to 10%, but standard French benchmarks didn't predict which model would do best.
The home languages of UK and US samples matter too. Only 78.4% of Londoners have English as their main language (91.1% across England and Wales). The next biggest languages nationally are Polish, Romanian, Panjabi and Urdu. Polish and Romanian are Ready. Panjabi and Urdu are Ready only on evidence from Indian audio (Sarvam and Gemini). ElevenLabs rates Urdu at 25 to 50% of words wrong, Google's Chirp 3 has no Pakistani Urdu locale, and we found no test on UK speakers. In the US, Spanish is spoken at home by 13.9% of people, followed by Chinese, Tagalog, Vietnamese and Arabic, all Ready, while Haitian Creole, with about a million speakers, is only Workable.
Markets we haven't mapped
Our 23-market map leaves out some markets that come up in almost every multi-country brief. Here is where their main languages stand.
| Market | Main language | Where it stands |
|---|---|---|
| France, Italy, Poland | French, Italian, Polish | Ready. ElevenLabs rates all three Excellent and Google has full support. |
| Vietnam | Vietnamese | Ready. Best real-world error about 10% (Google's Chirp 3). |
| South Korea | Korean | Ready. Best real-world error about 10% of characters (an Alibaba model); ElevenLabs 12%. |
| Japan | Japanese | Ready on clean audio, but real-world error was 28 to 44% of characters across Western providers. Part of that is scoring, since the same word can be written correctly in different scripts. Check a sample. |
| China | Mandarin | Ready. Regional varieties such as Wu and Xiang are much weaker, and mainly handled by Chinese providers. |
| Indonesia | Indonesian | Ready, but about 74% of people also use a regional language, and Javanese is only Workable and Sundanese Hard. |
| Saudi Arabia, UAE | Gulf Arabic | Workable. Best real-world error 17% on Saudi audio (Alibaba and Google's Chirp 3 within a point) and 26% on Emirati audio. |
| Egypt | Egyptian Arabic | Workable. Best real-world error 37% (an Alibaba model); the best Western provider, Gemini, 41%. |
| Australia | English | Ready. We found no accent-specific evidence, and more than 5.5 million people use another language at home. |
Code-switching
Real people don't speak one clean language at a time. A respondent in Nairobi slides between Swahili, English and Sheng in a single thought. In Johannesburg it's English threaded through isiZulu. In Lagos, English, Yoruba and Pidgin. In Manila, Taglish. In India, Hinglish.
Most speech models assume one language per sentence, and the gap shows. AfriSwitch, a 2026 benchmark of 61 hours of real conversation mixing English with 16 African languages and varieties, found no system below 24% on any of the 12 languages benchmarked so far. The best, Intron's own Sahara V2.5, averaged 36% across the 12 languages in its main results; Meta's Omnilingual, Google's Gemini and ElevenLabs averaged 51 to 56%. In the Philippines, where Taglish is everyday speech, the best real-world results for Tagalog were still around a quarter of words wrong.
Two things help. Use a speech model trained on mixed speech where one exists (Sarvam, for example, supports Hinglish). And don't assume an LLM can tidy it up afterwards: asking GPT-4o-mini to correct Hindi transcripts without examples took word error from 18% to 25%. Keep the original audio, and have a native speaker check the parts where people switch.
Recording conditions
Benchmark accuracy is measured on clean recordings. Your data is a voice note recorded on a cheap phone in a taxi or a market. Expect error rates roughly 1.5 to 2 times worse than the clean figures for single-language speech. On real Indian phone calls, the noisiest quarter of recordings scored about 1.65 to 1.75 times worse than the cleanest; on African conversational audio the gap ranged from 1.3 to 2.8 times.
The cause matters. WhatsApp's own audio compression (Opus) barely moves accuracy. Old-style 8 kHz phone lines do. And competing voices in the background do the most damage of all. So the cheapest improvement isn't a better model, it's asking people to record somewhere quiet, away from the TV and other conversations, and testing models on your own pilot audio rather than their published numbers.
So, AI or a team?
Put it together and it isn't “AI” or “humans”. It's a split that depends on the language, the provider and the people.
| Rating | How to set it up |
|---|---|
| Ready transcription, Strong translation | Run it with AI. Spot-check a sample, by region or community if your sample is mixed. |
| Ready or Workable transcription, Strong or Moderate translation | AI drafts, a native speaker checks the discussion guide before launch and the transcripts before analysis. |
| Hard transcription, but Strong or Moderate translation | Ask for typed answers, which AI can read and translate well, or have a person transcribe. |
| Basic translation | Keep a person in the loop for both transcription and translation. AI makes them faster, it doesn't replace them. |
Check the country, not just the language
Look at the home languages of the people you'll actually reach, and how many speak a Ready one.
Pick the engine per language
Shortlist two or three providers for each language and test them on your own pilot audio.
Set the variety
Choose the regional model where one exists, and budget a check by region or community.
Plan for mixing and noise
Ask for quiet recordings, keep the original audio, and have native speakers check the switched parts.
The right answer is AI-first, with people placed exactly where the technology is weak.
That's how research happens in people's own words at scale without quietly shipping bad data. “We can only do this in English” is no longer true for most large languages. Which languages, in which mode, with how much human review, is now a decision worth making on purpose.
Frequently asked questions
Can AI accurately transcribe African languages?
Some of them. Swahili and Afrikaans are the only two of the 44 sub-Saharan languages we rated that come out Ready. Most of the big ones, like isiZulu, Hausa, Yoruba and Amharic, are Workable: AI gives you a usable draft and a native speaker needs to check it. A long tail, including Tshivenda, isiNdebele and many smaller languages, still has no modern commercial speech-to-text. Mixing languages makes all of them harder.
Which AI transcription provider is best?
None of them across the board. On English the big providers all sit between about 2% and 5% of words wrong. Outside English the winner changes by language: on Indian phone calls Google's Gemini led on most languages but Amazon won Telugu, Kannada and Odia; on Cantonese Microsoft made about a quarter as many errors as Google's Chirp 3. Regional specialists such as Sarvam in India often beat all of them. Test two or three on your own audio before you commit.
Can Claude or ChatGPT transcribe interviews?
Claude can't. Its API takes text and images, not audio, so it isn't a transcription option, although it works well for translating and analysing a transcript once you have one. OpenAI does sell transcription models. Its GPT-4o Transcribe model was competitive on English but among the weakest on Indian phone audio and real-world Arabic and Asian audio in the independent tests we found. Its newer model hasn't been tested independently outside English yet.
Is English transcription solved?
On standard test sets, nearly. On real people, not quite. In the studies we found, error rates roughly doubled for Northern Irish and Northern English speakers, Black American speakers and Spanglish, and went up several times over for Scottish speech. If your sample spans regions or communities, check a few transcripts from each group, not just a few in total.
Why is translation usually easier than transcription?
Translation starts from text, and large language models have read far more text than speech models have heard audio. For about one language in five, the translation rating is a tier ahead of the transcription rating, and it's rarely behind. The catch is informal speech: in the 2025 machine translation shared task, transcribed speech was the hardest material to translate.
Should I let an LLM clean up a messy transcript?
Carefully. In a 2026 study, asking GPT-4o-mini to correct Hindi transcripts without examples pushed the word error rate from 18% to 25%. With good examples it helped in some setups and not others. Use a speech model trained on mixed speech where one exists, keep the original audio, and have a native speaker check the parts where people switch languages.
Do researchers still need human transcribers in 2026?
Yes, but in a smaller and more targeted role. For Ready languages a light spot-check is enough. For Workable languages a native speaker should check the discussion guide before launch and the transcripts before analysis. For Hard languages, ask for typed answers or have a person lead the transcription with AI helping.
That's our home ground.
We'll tell you which languages are ready to run on AI and which ones need a human hand, market by market.
Book a Demo →%202.png)


