Best Multilingual AI Voice Generators
Published September 2, 2026 · Every figure below was read from the vendor’s own pages on September 2, 2026. Where a vendor publishes an enumerated list, we counted it ourselves and say so. No listening test has been run, so nothing here says how any language sounds — this page is about coverage, not quality.
The short answer
Nine of the ten providers advertise a language count that is wrong against their own enumerated list. We counted all ten. Only Amazon Polly’s position survives, and only because it publishes no headline number to contradict.
Two vendors contradict themselves on a single page. ElevenLabs’ text-to-speech page says “over 70 languages” in its heading and “32+ languages” in its FAQ. Google’s product page says “75+ languages and variants” in its body while its own navigation, embedded in the same page, says “40+ languages”.
Most of the inflation is the languages-versus-locales gap, and almost nobody tells you which they are counting. Azure’s enumerated table holds 154 locales — but 82 base languages. Google’s holds 61 codes and 53 base languages. Cartesia is the only provider here that separates the two cleanly.
The second trap is availability. Google’s Gemini-TTS table lists 87 languages, of which 24 are generally available and 63 are preview. Speechify enumerates 6 fully supported languages against 25 in beta. A headline number that counts preview languages is not a number you can ship on.
So stop asking how many. Ask the three questions below about your specific language. The answer is checkable in minutes and the headline count never is.
Basis: vendor documentation only, checked September 2, 2026 — every count below is our own count of the vendor’s published list (how we verify).
On this page
Every count, recounted
We took each provider’s advertised figure, found its enumerated list, and counted the list. The third column is our count, not a vendor claim, and we say which page we counted.
| Provider | What the vendor advertises | Our count of its own list | What the gap is |
|---|---|---|---|
| Azure Speech | “100+” standard voices | 154 locales / 82 base languages | Counts locales; the figure understates locales and overstates languages |
| ElevenLabs | “over 70” and “32+” — both on the same page | 74 in docs / 76 in marketing | Two headline figures that cannot both be right, and two lists that disagree |
| Google Cloud | “75+ languages and variants” and “40+ languages” — same page | 61 codes / 53 base languages | Neither advertised figure matches the enumerated table |
| OpenAI | No total stated for text to speech | No TTS list published | The only enumerated list is Whisper’s, in a GitHub README, and it is speech recognition |
| Cartesia | “44 languages” | 44 codes in docs / 46 names in marketing | The docs match; the marketing list holds two more and drops Croatian and Gujarati |
| Amazon Polly | Table stated to be complete; no headline total | 42 entries | No gap — the only provider not advertising a number that conflicts with its list |
| MiniMax | “40” in docs, “32” on its audio product page | 40 entries | The list matches the higher figure; the product page contradicts both |
| Murf AI | “35 languages” and “35+” — same page | 39 locales / 33 base languages | 35 matches neither number derivable from its own library |
| Speechify | “30+ languages”; elsewhere “35 locales covering 30 distinct languages” | 31 live locale rows | Of those, only 6 are fully supported and 25 are beta |
| Resemble AI | “25 languages” and “20+ languages” — same page | 21 entries | Three figures for one model; and the vendor no longer sells voice generation |
Amazon Polly is the only provider on this page whose published position survives a recount intact — because it does not lead with a headline number. It publishes a table it states is complete, we counted 42 entries, and there is nothing to reconcile. That is not a small thing. It is the difference between a figure you can plan around and a figure you cannot.
Two notes on reading that table. Our counts are of what each vendor publishes, so a larger number is not better: Azure’s 154 is locales, Amazon’s 42 is language variants, and they are not the same unit. And a count is not a capability — none of these numbers tells you whether a given language is any good, which is a question no page on this site can answer until our benchmark runs.
The three questions that replace “how many”
Since no headline count is reliable, here is what to establish instead. All three are answerable from vendor documentation in a few minutes, for your language specifically.
- Is my language in the enumerated list? Not in the marketing count — in the table. Six of the ten publish a readable enumerated list. If your language is not on it, the headline number is irrelevant to you.
- Is it generally available, or preview? This is the question that costs people money. Google’s Gemini-TTS table marks 63 of its 87 languages preview. Speechify marks 25 of its 31 beta. Preview support can change or disappear without the notice a GA service would give you.
- Is the vendor counting languages or locales? If you need Brazilian and European Portuguese as distinct voices, a “languages” count of 44 may deliver one of them. Cartesia is the only provider here that answers this cleanly: it counts languages and treats locales as a separate axis with its own API field — and its separate locale table covers just 5 of its 44 languages.
A fourth question applies if you use an API: does coverage change by model? It does for at least three providers, and the headline number is usually attached to none of them.
Languages against locales
This is where most of the apparent difference between vendors evaporates. A locale is a language plus a region — pt-BR and pt-PT are two locales of one language. Counting locales roughly doubles a catalog without adding a language.
| Provider | Locales / codes | Distinct base languages |
|---|---|---|
| Azure Speech | 154 | 82 |
| Google Cloud | 61 | 53 |
| Murf AI | 39 | 33 |
Azure is the clearest case: its enumerated text-to-speech voices table holds 767 voice rows across 154 distinct locales, which reduce to 82 distinct base language subtags. Its own overview advertises “100+”, a figure that sits between the two and matches neither. Azure’s documentation is at least explicit that it presents coverage in locales rather than languages, which is more than most.
Murf’s locale codes are not standard in places, which makes its count non-comparable with the BCP-47 counts every other vendor uses. If you are building a language picker against several providers, that is a real integration cost and it is not mentioned anywhere on Murf’s marketing pages.
Generally available against preview
The second trap, and the more expensive one. A language marked preview or beta is not a language you should build a launch around, and headline counts fold the two together silently.
- Google Cloud, Gemini-TTS. The docs table carries a “Launch readiness” column. We counted 87 rows: 24 marked GA, 63 marked Preview. Google’s marketing describes the same models as available “in 75+ locales” with no GA qualifier anywhere near it. If you take the marketing figure at face value, roughly three quarters of what you are counting is preview.
- Speechify. Its language-support page splits into three tables: Fully Supported, Beta, and Coming Soon. We counted 6 fully supported, 25 beta, and 17 not yet available. Its headline is “30+ languages”. Its streaming-native
simba-3.0model officially supports exactly seven locales. - Google Cloud, standard voices. The docs overview table gives a launch stage per voice type; one Studio configuration is listed as Experimental while every other row is GA. That is a well-documented position and worth crediting.
- Cartesia. Odia and Urdu are identified as new in Sonic 3.6 and are not marked preview or beta — the vendor ships them as available. Whether that is confidence or an omission, its documentation takes a position.
One inconsistency worth knowing if you are evaluating Speechify: Italian appears in its beta table and simultaneously among the seven locales simba-3.0 “officially supports”, on the same page. The page does not say whether Italian is beta or generally available overall.
Who publishes a list you can actually read
A count you cannot check is not evidence. Six providers publish an enumerated list a buyer can read and count; four do not, for four different reasons.
| Provider | Readable list? | Detail |
|---|---|---|
| Amazon Polly | Yes | Numbered table Amazon states is complete, with a column per voice engine |
| Azure Speech | Yes | 767 voice rows with a Type column; the page warns coverage is feature-dependent |
| Google Cloud | Yes | 2,032 voice rows; separate Gemini-TTS table with a launch-readiness column |
| Cartesia | Yes | 44 language codes in a dated-snapshot table, plus a separate locale table |
| MiniMax | Yes | A self-numbering table of exactly 40 entries, identical on both platform guides |
| Speechify | Yes | Three tables split by support status — the clearest availability disclosure here |
| ElevenLabs | Partly | Two lists that disagree: 74 entries in docs, 76 names in marketing |
| Murf AI | Partly | Gen2 locale lists are truncated in both the rendered page and the markdown, so that section cannot be fully counted; the docs’ “Supported Locales” section is prose only and the full list needs an authenticated API call |
| Resemble AI | Partly | 21 names on a model page; the API documentation index contains zero occurrences of “language”, “locale” or “multiling” |
| OpenAI | No | No text-to-speech language list exists — see below |
Speechify deserves a specific credit here that its headline number does not earn it: splitting its languages into Fully Supported, Beta and Coming Soon is the most useful availability disclosure in this comparison. It is also what makes its “30+” claim look generous, since only six are fully supported. Both things are true, and a vendor that publishes the data to undercut its own marketing is doing the reader a favor.
Coverage changes by model
Most of these providers sell several models, and language coverage is not the same across them. The advertised number is usually attached to no model at all.
- Murf states it directly — “Supported Voices & Languages vary by model— Gen2 and Falcon 2 each have their own availability” — and then publishes a single undifferentiated “35” with no model attached. Its current model, Falcon 2, enumerates 28 locales untruncated. The deprecated Gen2 enumerates at least 35, from a truncated list.
- Google runs entirely separate language tables for Gemini-TTS and for its other voice types, with different availability rules.
- Speechify’s
simba-3.0covers seven locales whilesimba-multilingualis described as “35 locales covering 30 distinct languages” — a figure that cannot be derived from the tables on its own language-support page. - Azure’s enumerated table has a Type column taking five values whose locale coverage is very uneven.
If you are choosing on languages and you use the API, the number that matters is the one attached to the model you will call. For Murf that is 28, not 35.
The provider that publishes no list at all
OpenAI is the only provider here with no text-to-speech language list of any kind, and the way that gap presents is worth describing because it is easy to miss.
Its text-to-speech guide does carry an enumerated list of 57 language names — but that list is presented as Whisper’s, and Whisper is speech recognition. It describes what OpenAI can transcribe, not what it can speak, and the authoritative version of it lives in a GitHub README rather than the API documentation. A separate 16-language enumeration exists for custom-voice consent recordings, which is a different thing again.
There is also a documentation defect a developer will hit: the guide’s own worked example posts a language field to the speech endpoint that the endpoint’s API reference does not document. The reference lists input, model, voice, instructions, response_format, speed and stream_format — and a case-insensitive search of that page for “language” returns nothing. The same example also sends format where the reference documents response_format.
So if you are choosing OpenAI for multilingual work, you are choosing without a coverage list, on a parameter the reference does not acknowledge. That may still be the right choice for other reasons — our OpenAI review sets them out — but it is not a choice you can make on documented language support.
How to choose for your languages
- Write down your actual list, as locales, not languages —
pt-BR, not “Portuguese”. This single step resolves most of the confusion on this page. - Check each provider’s enumerated table for those exact codes. Ignore the headline number entirely; it has failed against its own list for every provider we checked.
- Check the availability column where one exists. If your language is preview or beta, treat it as something that may change without a GA deprecation notice.
- Check the model. If you will call one specific model, find that model’s coverage rather than the vendor’s aggregate.
- Prefer a provider that enumerates over one that advertises. Amazon Polly publishes a table it says is complete and no conflicting headline; Speechify publishes its own beta split. Those are the two positions on this page a buyer can plan against.
- Then test the languages you actually need, because coverage is not quality, and nothing on this page — or on any page of this site today — tells you how a given language sounds.
Frequently asked questions
Which AI voice generator supports the most languages?
By our count of published lists, Azure Speech enumerates the most at 154 locales, which reduce to 82 distinct base languages. But the counts are in different units across vendors, so comparing them directly is not meaningful. The better question is whether your specific language is enumerated and generally available.
Why does every vendor’s number seem wrong?
Because most of them count locales while advertising languages, fold preview languages into the headline, and attach one figure to several models with different coverage. Two — ElevenLabs and Google — publish two incompatible figures on a single page.
What is the difference between a language and a locale?
A locale is a language plus a region: pt-BR and pt-PT are two locales of Portuguese. Counting locales roughly doubles a catalog without adding a language. Azure’s 154 locales are 82 base languages; Google’s 61 codes are 53.
Is a preview language safe to build on?
That is a risk decision, not a documentation one, and the vendors do not make it for you. What we can tell you is the scale: 63 of the 87 languages in Google’s Gemini-TTS table are marked preview, and 25 of Speechify’s 31 live locales are beta. Neither headline count mentions it.
Which provider has the most trustworthy language documentation?
Amazon Polly, for not advertising a number that conflicts with its own table, and Speechify, for publishing the beta split that undercuts its own headline. Cartesia is the only one that cleanly separates languages from locales.
Does OpenAI support multiple languages for text to speech?
It publishes no text-to-speech language list. The 57-language list on its guide is Whisper’s speech-recognition coverage, and its speech endpoint’s reference does not document the language parameter its own example sends.
Do these numbers tell me anything about quality?
No. Coverage and quality are different questions, and this page answers only the first. We have run no listening test in any language, so we cannot tell you whether a supported language is well supported.
Sources and what we could not verify
Every advertised figure and every enumerated list was read from the vendor’s own pages on September 2, 2026. Every count in the third column of the first table is ours, made from the vendor’s published list, and the page it came from is named in the row. Per-provider detail is in the ten reviews linked throughout, and the method is on the methodology page.
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 2, 2026:
- How any of these languages sound. No listening test has been run. Coverage is not quality, and a language being enumerated says nothing about whether its output is usable.
- Which of each vendor’s conflicting figures is correct. Where a vendor publishes two, we publish both and neither is endorsed.
- Murf’s complete Gen2 language list. The per-voice locale lists are truncated in the rendered page and stay truncated in the markdown, and the docs’ “Supported Locales” section is prose only — the full list is reachable only through an authenticated API call, which we did not make.
- Whether Speechify’s Italian is beta or generally available. It appears in both the beta table and the officially-supported list on the same page.
- OpenAI’s text-to-speech language coverage. No list is published, and the
languageparameter its guide uses is undocumented in the endpoint reference. - Whether MiniMax’s 40 are languages or locales. No MiniMax page we read states which.
- Resemble AI’s position. Its own documentation lists all previous text-to-speech models as end of life while its marketing site presents Chatterbox Multilingual V3 as current. Separately, the company no longer sells voice generation.
- Whether any of these lists is current. Only Cartesia publishes its language list in a dated snapshot. The rest carry no version and no effective date.
Change log
September 2, 2026 — First publication. All figures read from official vendor documentation on September 2, 2026; all list counts are our own and identified as such.
Published September 2, 2026 · This page shows no audio and reports no listening results, because our benchmark has not run and does not yet cover multilingual output. Language lists change without notice and only one vendor dates its own; corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.