You have research to run in Swahili, or isiZulu, or Hausa, or across a dozen markets at once, and you want one answer: can AI transcribe and translate it, or do you still need a room full of people doing it by hand? Here is the honest, language-by-language answer, in tables built to scan.
For most of the languages you care about, AI now does the heavy lifting and a small human layer catches what it misses. For a stubborn minority, that ratio flips and people still need to lead. The whole game is knowing which bucket your language falls into before you start, which is exactly what the table below is for.
The language table
Two separate ratings, because they're two separate jobs. Transcription turns a voice note into text in the same language. Translation turns that text into English. Transcription is the harder job and where nearly all the difficulty lives; translation runs roughly a tier ahead almost everywhere.
● Ready AI does the job, light spot-check only. ◐ Workable AI drafts, a human closes the gap. ○ Hard Human-led, AI assists.
Grouped by region, with the languages most relevant to emerging-market fieldwork first. "Approach" is the setup we'd actually reach for.
| Language | Speakers / region | Transcription | Translation | Recommended approach |
|---|---|---|---|---|
| Southern Africa | ||||
| Afrikaans | 12M · ZA, NA | ● Ready | ● Strong | Scribe or Chirp 3, light spot-check. |
| isiZulu | 28M · ZA | ◐ Workable | ◐ Moderate | Chirp 3 (streaming) or Scribe (batch), human QA on nuance. |
| isiXhosa | 19M · ZA | ◐ Workable | ◐ Moderate | Chirp 3, human QA on nuance and names. |
| Sesotho | 14M · ZA, LS | ◐ Workable | ◐ Moderate | Chirp 3 is the only real option currently, human QA. |
| Sepedi (N. Sotho) | 14M · ZA | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Setswana | 14M · ZA, BW | ◐ Workable | ◐ Moderate | Chirp 3, human QA. |
| Shona | 15M · ZW | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Chichewa (Nyanja) | 20M · MW, ZM | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Xitsonga | 3.5M · ZA | ○ Hard | ○ Basic | Meta MMS (self-hosted) or a human transcriber. |
| siSwati | 3M · ZA, SZ | ○ Hard | ○ Basic | Meta MMS or a human transcriber. |
| Tshivenda | 2M · ZA | ○ Hard | ○ Basic | Meta MMS or a human transcriber. |
| isiNdebele | 1.5M · ZA | ○ Hard | ○ Basic | Meta MMS or a human transcriber. |
| East Africa | ||||
| Swahili | 95M · KE, TZ, UG, CD | ● Ready | ● Strong | Scribe (best accuracy) or Chirp 3, light spot-check. |
| Amharic | 78M · ET | ◐ Workable | ◐ Moderate | Chirp 3 or AWS, human QA. |
| Somali | 22M · SO, ET, KE | ◐ Workable | ◐ Moderate | Chirp 3 or AWS, human QA. |
| Kinyarwanda | 15M · RW | ◐ Workable | ◐ Moderate | Chirp 3 with human QA, or fine-tune for scale. |
| Luganda | 10M · UG | ○ Hard | ◐ Moderate | Fine-tuned Whisper (Sunbird), human-led otherwise. |
| Oromo | 46M · ET, KE | ◐ Workable | ◐ Moderate | Chirp 3 (Afaan Oromoo), human QA. |
| Tigrinya | 9M · ER, ET | ○ Hard | ○ Basic | Chirp 3 partial coverage, else Meta MMS, human-led. |
| Kirundi | 12M · BI | ○ Hard | ○ Basic | Meta MMS, human-led. |
| West & Central Africa | ||||
| Hausa | 94M · NG, NE, GH | ● Ready | ● Strong | Scribe (best) or Chirp 3, light spot-check. |
| Yoruba | ~50M · NG, BJ | ◐ Workable | ◐ Moderate | Chirp 3 or fine-tuned Whisper, human QA on tone. |
| Igbo | ~30M · NG | ◐ Workable | ◐ Moderate | Chirp 3 is the only commercial option, human QA. |
| Lingala | 40M · CD, CG | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Akan (Twi) | 20M · GH | ◐ Workable | ◐ Moderate | Chirp 3, human QA. |
| Wolof | 12M · SN | ○ Hard | ◐ Moderate | Scribe or Chirp 3, patchy, human-led. |
| Fula (Fulah) | 40M · W. Africa | ○ Hard | ◐ Moderate | Scribe or Meta MMS, human-led. |
| Nigerian Pidgin | 60–120M · NG | ○ Hard | ◐ Moderate | No dedicated speech-to-text support yet, human transcriber then LLM cleanup. |
| Mossi (Moore) | 9M · BF | ○ Hard | ○ Basic | Meta MMS, human-led. |
| Kanuri | 9M · NG, NE, TD | ○ Hard | ○ Basic | Meta MMS, human-led. |
| Middle East & Arabic | ||||
| Arabic (MSA) | 335M · MENA | ● Ready | ● Strong | AWS or Chirp 3, light spot-check. |
| Egyptian Arabic | 118M · EG | ● Ready | ● Strong | AWS dialect locale, light spot-check. |
| Levantine Arabic | 58M · SY, LB, JO | ◐ Workable | ◐ Moderate | AWS or Chirp 3, human QA on dialect. |
| Moroccan Arabic (Darija) | 21M · MA | ◐ Workable | ◐ Moderate | AWS or Chirp 3, human QA (distinct dialect). |
| Sudanese Arabic | 54M · SD | ◐ Workable | ◐ Moderate | Scribe or Azure, human QA. |
| Persian (Farsi) | 82M · IR | ● Ready | ● Strong | Scribe or Chirp 3, light spot-check. |
| Kurdish (Kurmanji) | 26M · TR, IQ, SY | ◐ Workable | ◐ Moderate | Chirp 3 or AWS, human QA. |
| Hebrew | 9M · IL | ● Ready | ● Strong | Scribe or Chirp 3, light spot-check. |
| South Asia | ||||
| Hindi | 611M · IN | ● Ready | ● Strong | Any major provider, light spot-check. |
| Bengali | 274M · BD, IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Urdu | 246M · PK, IN | ● Ready | ● Strong | Scribe or Chirp 3, light spot-check. |
| Tamil | 86M · IN, LK | ● Ready | ● Strong | Any major provider, light spot-check. |
| Telugu | 96M · IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Marathi | 99M · IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Gujarati | 62M · IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Kannada | 59M · IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Malayalam | 38M · IN | ● Ready | ● Strong | Scribe or Google, light spot-check. |
| Punjabi | 90M · PK, IN | ◐ Workable | ● Strong | Chirp 3 or Scribe, human QA. |
| Odia (Oriya) | 38M · IN | ◐ Workable | ◐ Moderate | Google or AWS, human QA. |
| Sindhi | 37M · PK, IN | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Assamese | 24M · IN | ◐ Workable | ◐ Moderate | Scribe or Google, human QA. |
| Nepali | 32M · NP | ◐ Workable | ● Strong | Scribe or Chirp 3, human QA. |
| Sinhala | 26M · LK | ◐ Workable | ◐ Moderate | Scribe or Chirp 3, human QA. |
| Bhojpuri | 53M · IN | ○ Hard | ◐ Moderate | Meta MMS, human-led. |
| East & Southeast Asia | ||||
| Mandarin Chinese | 1,184M · CN | ● Ready | ● Strong | Any major provider, light spot-check. |
| Japanese | 126M · JP | ● Ready | ● Strong | Any major provider, light spot-check. |
| Korean | 82M · KR | ● Ready | ● Strong | Any major provider, light spot-check. |
| Indonesian | 255M · ID | ● Ready | ● Strong | Any major provider, light spot-check. |
| Vietnamese | 97M · VN | ● Ready | ● Strong | Any major provider, light spot-check. |
| Thai | 71M · TH | ● Ready | ● Strong | Scribe, Chirp 3, or Deepgram, light spot-check. |
| Tagalog | 87M · PH | ● Ready | ● Strong | Any major provider, light spot-check. |
| Cantonese (Yue) | 86M · HK, CN | ◐ Workable | ● Strong | Chirp 3 or Scribe, human QA. |
| Cebuano | 28M · PH | ◐ Workable | ● Strong | Chirp 3 or Scribe, human QA. |
| Javanese | 69M · ID | ◐ Workable | ◐ Moderate | Chirp 3 or Scribe, human QA. |
| Burmese | 43M · MM | ◐ Workable | ◐ Moderate | Chirp 3 or Scribe, human QA. |
| Khmer | 18M · KH | ◐ Workable | ◐ Moderate | Chirp 3 or Scribe, human QA. |
| Europe & global majors | ||||
| English | ~1.5B · global | ● Ready | ● Strong | Any provider, Deepgram is strong here too. |
| Spanish | 561M · ES, LatAm | ● Ready | ● Strong | Any provider. |
| French | 334M · FR, Africa, CA | ● Ready | ● Strong | Any provider. |
| Portuguese | 269M · BR, PT | ● Ready | ● Strong | Any provider. |
| Russian | 210M · RU, CIS | ● Ready | ● Strong | Any provider. |
| German | 133M · DE, AT, CH | ● Ready | ● Strong | Any provider. |
| Italian | 66M · IT | ● Ready | ● Strong | Any provider. |
| Polish, Dutch, Ukrainian & other European | Europe | ● Ready | ● Strong | Any provider, all well covered. |
One thing about these ratings: they move. Investment in African and emerging-market languages has accelerated sharply, so a language sitting in "Hard" today can be "Workable" within a year. The floor is rising fast, and faster for the languages with the most speakers.
Which model for which job
"Use AI" isn't one decision. The engines differ enormously, and routing to the wrong one is the most common reason a multilingual study underperforms. A serious operation sends each language to its best model rather than forcing everything through a single provider, which is the logic behind the "recommended approach" column above.
| Model | Best for | Strengths | Where it falls short |
|---|---|---|---|
| ElevenLabs Scribe | Best raw accuracy on covered languages | Tops several 2025-2026 accuracy benchmarks (96.7% on English), including a number of African languages. Batch and real-time modes. | Benchmarks lean on clean, read-aloud audio, so real voice notes land softer. |
| Google Chirp 3 | Broad African & emerging-market coverage | One of the widest commercial footprints for African languages, streams live, with denoising and speaker separation built in. | Rarely publishes per-language accuracy figures, so test on your own audio before committing. |
| OpenAI Whisper / gpt-4o-transcribe | Fine-tuning to a bespoke edge case | 99 languages on paper, open weights that fine-tune from unusable to strong with 50-200 hours of local audio. | OpenAI itself flags 20 of the 99 as having no training data. Weak zero-shot on many African and tonal languages; built-in translation outputs English only. |
| Deepgram | English, speed, low cost | Fast and inexpensive, strong on English and several dozen European and Asian languages. | No African-language coverage as of writing. Not a candidate for African local-language work. |
| AWS Transcribe / Azure | Enterprise breadth on major languages | Wide coverage of major world languages, strong on Arabic dialects, well-supported in-cloud. | Thins out on the tail. Patchy on smaller African languages. |
| Meta MMS (open source) | The long tail no one else covers | Over 1,100 languages for speech recognition, including minority languages no commercial API touches. The answer for the "Hard" tier. | Open source, so you need the engineering resource to host and run it yourself. |
No single model wins everywhere, and the same is true for translation. Google Translate has the widest coverage, Meta's open NLLB is strong on low-resource languages, and modern general-purpose LLMs now match or beat dedicated translation engines on the top languages and handle slang and mixed text more gracefully, while still weakening on the minority tail.
Code-switching is the thing that breaks everything
Here's the failure mode that catches people out, because it shows up in no benchmark. Real people don't speak one clean language at a time. A respondent in Nairobi slides between Swahili, English, and Sheng within a single thought. In Johannesburg it's English threaded through isiZulu. In Lagos, English, Yoruba, and Pidgin at once. In India, Hinglish and Tanglish.
Almost every speech model is trained on the assumption of one language per utterance, so code-switching is exactly where they quietly fall apart, and in emerging markets this isn't an edge case. For a large share of urban respondents, mixing languages is simply how they talk.
The silver lining: large language models handle code-switched text far better than speech engines handle code-switched audio, since they're trained on the messy way people actually write. The winning pattern is a decent transcript first, then an LLM to clean, interpret, and translate the mix, rather than expecting the speech engine to nail it in a single pass. Even so, this remains the single biggest reason to keep a human in the loop on the harder tiers. Code-switching is where the machine is most confidently wrong.
The recording environment matters more than the model
Benchmark accuracy gets measured on clean studio audio. Your data is a voice note recorded on a cheap phone, in a taxi, in a market, with wind and a TV in the background, then squeezed through WhatsApp compression.
Based on our own fieldwork, expect error rates roughly 1.5 to 2 times worse than any benchmark number, purely from the audio conditions. That's not the model failing, it's physics. Two things follow. Denoising before transcription genuinely moves the numbers, which is why models with built-in cleanup have a quiet advantage on field audio. And a little guidance to respondents, find a quiet spot, hold the phone close, improves data quality more cheaply than any model upgrade. The environment is the variable people ignore and then blame the AI for.
So, AI or a team?
Put it together and the answer isn't "AI" or "humans." It's a ratio that shifts with the job.
| Your study | The right setup |
|---|---|
| Ready-tier | AI does the work, a light review pass is enough. A full manual team here is money set on fire. |
| Workable-tier | AI drafts, a human closes the gap on nuance and key findings. The sweet spot for serious work, faster and cheaper than fully manual without losing trust. |
| Hard-tier, heavy code-switch, or rough audio | AI assists a human rather than replacing them. People still need to lead, but AI makes them faster. |
The old model, a large human team transcribing and translating everything by hand, is now the wrong answer for almost every study. It's slow, it doesn't scale, and it's expensive where it doesn't need to be. But pure AI with no human anywhere is equally wrong for the hard cases, because it fails silently, and you find out when a client questions a finding you can't defend.
The right answer is AI-first, with humans placed exactly where the technology is weak.
Route each language to its best model, clean the audio, let AI carry the volume, and spend human hours on the specific failure points: the code-switching, the low-resource languages, the nuance that carries the insight. That's how research happens in people's own words, at scale, without burning budget or quietly shipping bad data. "We can only do this in English" is no longer a real constraint. Which languages, in which mode, with how much human review, is now a strategy decision, and one worth making on purpose.
Frequently asked questions
Can AI accurately transcribe African languages?
For high-resource languages like Swahili, Hausa, and Afrikaans, yes, largely. For most other African languages, AI produces a usable draft that still needs a human reviewer, and for a small number of low-resource languages, human-led transcription remains the safer default.
Why is translation usually easier than transcription?
Translation starts from clean text, while transcription has to first turn messy, accented, background-noise-laden audio into that text. Nearly all of the difficulty in a multilingual pipeline lives in getting an accurate transcript, not in the translation step that follows.
Which AI transcription model is best for research?
There's no single winner. ElevenLabs Scribe tends to lead on raw accuracy for covered languages, Google Chirp 3 has the broadest African-language reach, and Meta's open-source MMS model is the fallback for the long tail nothing commercial supports. Serious operations route each language to whichever model handles it best.
Why does code-switching break AI transcription?
Most speech models assume one language per utterance. Real conversations, especially in multilingual urban settings, blend two or three languages in a single sentence, which is exactly the pattern these models weren't trained to expect.
Do researchers still need human transcribers in 2026?
Yes, but in a much smaller role than before. Human effort is best spent reviewing AI drafts for nuance on workable-tier languages and leading transcription outright on hard-tier languages, rather than transcribing everything from scratch.
That's our home ground.
We'll tell you honestly which languages are ready to run on AI, and which ones need a human hand, market by market.
Book a Demo →%202.png)
.png)

