Real-Time AI Voice Generators
Published September 2, 2026 · Every figure below was read from the vendor’s own pages on September 2, 2026. This site has run no latency measurement. Every number on this page is a vendor claim, labelled as one, and no table here is sorted by speed.
The short answer
We will not tell you which of these is fastest, and neither can anyone else from published figures. The numbers vendors publish measure different things, and most vendors never say which thing.
One vendor publishes three different latency numbers for the same model — and they need not conflict, because each names a different metric. ElevenLabs states ~75 ms, 100–150 ms and sub-500 ms for Flash v2.5 in three places.
That same vendor tells you the problem in its own documentation: “Latency in audio generation is deceptively simple-sounding but involves several distinct phenomena that are easy to conflate.” It then says its own headline figure is not the one that matters.
What is comparable is agent capability, and it is checkable. Does barge-in exist as a configurable parameter? Is turn detection parameterised? Is telephony documented? Seven providers answer yes, two sell no agent product at all, and one markets an agent with no developer documentation.
Basis: vendor documentation only, checked September 2, 2026 — no measurements, no accounts, no API calls (how we verify).
On this page
Why this page publishes no speed ranking
Ranking latency requires measuring it: the same input, the same transport, the same region, the same concurrency, repeated enough times to report a distribution rather than a best case. We have not done that, and until we do a ranking from this site would be an ordering of other people’s marketing.
It would also be wrong in a specific and demonstrable way. The published figures are not measurements of the same quantity. Putting ElevenLabs’ 75 ms next to Cartesia’s 90 ms produces an ordering, and the ordering is an artifact of which metric each vendor chose to publish rather than of how fast anything is. The next section shows that from the vendors’ own pages.
So this page does two things instead. It sets out what each vendor actually claims and what that claim measures, so you can stop treating those numbers as a league table. And it compares the things that are documented and checkable — interruption handling, turn detection, telephony and billing — which turn out to matter more for a working voice agent than any millisecond figure.
The metric problem
Below is every latency figure we found, with the metric the vendor named and whether the vendor defined it. The table is ordered alphabetically, not by number, and the numbers are not comparable with each other. That is the point of the table.
| Provider | Vendor’s figure | Metric named | Defined by the vendor? |
|---|---|---|---|
| Cartesia | “sub-90ms latency” | None — the bare word “latency” | No. A separate page instead describes streaming “the first byte of audio in about 90ms” |
| ElevenLabs | “~75ms†” | “model inference latency” | Yes — and the vendor says it is not the number that matters |
| ElevenLabs | 100–150 ms / 150–200 ms by region | “TTFB” | No — named but never expanded or defined on that page |
| ElevenLabs | “sub-500ms” | “end-to-end latency” | No — named, not defined on that page |
| Google Cloud | No number published | — | — |
| Murf AI | “55ms” | “model inference” and “model latency”, used interchangeably | No |
| Murf AI | “~100ms” (docs) / “130ms” (marketing) | “time-to-first-audio” | Yes, on one page: “the delay between a synthesis request and the first audio frame” |
| OpenAI | No number published | “Speed” — an ordinal rating, “Fast” or “Medium” | Not applicable — no time value is published at all |
| Speechify | “<300ms” | “TTFB” and “startup latency”, on different surfaces | No — neither term is expanded or defined anywhere we read |
Read the second, third and fourth rows together. They are all ElevenLabs, and they are all about the same model family. 75 ms is model inference. 100–150 ms is TTFB for North America, Europe and South East Asia, rising to 150–200 ms for South and North East Asia. Sub-500 ms is end-to-end. Three numbers, three metrics, one product, and no contradiction between them — only one is the number a user would experience. We have verified none of the three; they are the vendor’s.
Cartesia is the opposite case. Its published figure carries no metric at all — “sub-90ms latency” — and its documentation landing page separately describes Sonic 3.5 as streaming “the first byte of audio in about 90ms”. Those are different quantities sharing a number. Meanwhile Sonic 3.6, which Cartesia announced as generally available, carries no latency figure at all: its page contains zero values of the form “nms”. The headline number belongs to the previous model.
Murf publishes the widest spread and the clearest definition. Its marketing pairs “55ms model inference” with “130ms end-to-end latency” on one page, its documentation says “~100ms time-to-first-audio” for Falcon 2, and one Murf page defines time-to-first-audio properly as the delay between a synthesis request and the first audio frame. It is the only vendor here besides ElevenLabs to define anything.
And Murf defines the same metric two different ways. Its latency documentation names the metric “Time to First Audio Byte” and operationalises it against a concrete curl field, time_starttransfer. Its benchmark page defines time-to-first-audio as the delay between a synthesis request and the first audio frame. A byte and a frame are not the same event, so two Murf pages using the same acronym are measuring two different things. That is the metric problem happening inside a single vendor.
Three vendors publish no number at all. Google Cloud attaches no latency figure to text-to-speech anywhere we read, and neither does MiniMax on either of its platforms. OpenAI publishes an ordinal instead — a “Speed” label reading “Fast” or “Medium”. Neither absence is a weakness; a vendor that publishes no number is not making a claim you have to discount.
The vendor that explains the problem itself
The most useful document in this entire subject is published by ElevenLabs, and it undercuts ElevenLabs’ own marketing. It is worth quoting at length because no summary improves on it.
The page opens: “Latency in audio generation is deceptively simple-sounding but involves several distinct phenomena that are easy to conflate.”
It then defines the two quantities that get conflated. First: “Model inference latency is the time the model spends generating audio. ElevenLabs Flash models achieve ~75ms model inference for typical short inputs. This is an internal measurement, excluding network round-trips and application overhead.”
And second: “Time-to-first-audio (TTFA) is the elapsed time from when your application initiates a request to when the first audio sample actually plays for the end user. This is almost always the number that matters for user experience, and it is always larger — often substantially larger — than model inference latency alone.”
Elsewhere the vendor is equally direct about its headline figure: “75ms refers to model inference time only. Actual end-to-end latency will vary with factors such as your location & endpoint type used.” And of that same figure: “It is a useful reference point for comparing models, not a guarantee for every request.”
So the vendor that publishes the lowest headline number in this market also publishes, in its own documentation, the statement that the number is not what a user experiences and that the figure which is — TTFA — is always larger. No TTFA figure is attached to it anywhere on that page.
The dagger on “~75ms†” resolves to a scope note rather than a definition: “† Excluding application & network latency”. Two of the three places its ~280 ms figure for the conversational model appears carry no dagger at all, so the exclusion travels with the number only some of the time.
None of this makes ElevenLabs’ claims dishonest. It makes them precise in a way the rest of the market’s are not, and it means the one vendor documenting the problem is the one whose headline number gets quoted most often without it.
What no vendor publishes
Across every page we read, for all ten providers, we found none of the following attached to any latency figure:
- A percentile. No p50, p95 or p99 anywhere, with one exception: Cartesia describes an agent figure as a “median latency” without defining the underlying metric. Every other number is presented without a distribution, so it is impossible to know whether it is a typical case or a best case.
- A measurement date. Not one latency claim carries the date it was measured. Prices on these sites carry dates; speed claims do not.
- An input length, with one exception: ElevenLabs conditions its 75 ms on “typical short inputs” in one place. Since synthesis time scales with text, a figure without an input length is not reproducible.
- A concurrency condition. Murf asserts its figure holds “even with 10,000 concurrent calls”, which is a claim rather than a condition. Nobody else states what load their number was taken under — and ElevenLabs separately documents that exceeding your concurrency limit queues requests, adding “~50ms of latency”.
- A network path or client location, except in ElevenLabs’ regional TTFB table, which is the only region-conditioned latency data published by any provider here.
A figure with no percentile, no date, no payload, no load and no location is not something you can plan a system around. It is a marketing number, and treating it as anything else is the mistake this page exists to prevent.
What is comparable: agent capability
Here is the constructive half. Whether a provider documents interruption handling as a configurable parameter is not a matter of opinion or measurement — it either appears in the API reference or it does not. The same goes for turn detection and telephony. These are the things that decide whether a voice agent works, and unlike latency they are checkable today.
| Provider | Agent product | Interruption | Turn detection | Telephony |
|---|---|---|---|---|
| Amazon | Amazon Lex V2, which speaks using Polly voices | allowInterrupt, on by default, at prompt and per-attempt level | Two required millisecond parameters | Documented |
| Azure Speech | Voice Live API | interrupt_response, plus a published filler-word filter to reduce false triggers | Four named types with published defaults and ranges | Documented |
| Cartesia | Managed Agents / Line | Asserted in prose; not a configurable field in the agent configuration reference | Parameterised on the standalone speech-to-text endpoint, not on the agent product | Own SIP endpoint, bring-your-own-carrier |
| ElevenLabs | ElevenAgents | “Interruption handling”, configurable on/off | “Take turn after silence” plus “Turn eagerness” with three modes | SIP trunking, inbound and outbound |
| Google Cloud | Conversational Agents (separate billed product) | “Barge-in”, a per-level toggle | End-of-speech sensitivity, smart endpointing, no-speech timeout | Hosts phone numbers directly |
| MiniMax | None | — | — | — |
| Murf AI | Marketed, no developer documentation | Marketing prose only — no term of art, no parameter | Described as customisable; no parameters published | Not documented |
| OpenAI | Realtime API, plus RealtimeAgent in the Agents SDK | “barge-in”, via turn_detection.interrupt_response | server_vad and semantic_vad — silence-based and classifier-based | Own SIP endpoint; does not sell phone numbers |
| Resemble AI | None | — | — | — |
| Speechify | SpeechifyAI Agents (beta) | interruption_sensitivity, a three-level enum with published thresholds | turn_handling object with published default wait windows | Phone numbers via LiveKit, Twilio or bring-your-own SIP |
The gap between marketing and documentation is where this table earns its place. Murf markets AI voice agents; its API documentation index contains no agent endpoints and no conversational API. Cartesia’s documentation says its agent “handles turn-taking and interruptions” in prose, but interruption is not exposed as a configurable field in the agent configuration reference, and its rich turn-detection parameters belong to a different product. Both may work well. Neither publishes what you would need to tune it.
Two vendors sell no agent product at all. MiniMax documents text-to-speech, async long-form, cloning, voice design and voice management — and no realtime speech-to-speech surface. The word “Agent” appears throughout its documentation but never in a voice sense. Resemble AI likewise, which is unsurprising given it no longer sells voice generation.
The most detailed interruption controls belong to the platform vendors and OpenAI. Azure publishes an English filler-word list whose stated purpose is reducing false barge-in triggers — a level of specificity nobody else reaches. OpenAI documents two turn-detection modes, one silence-based and one classifier-based, and is explicit that on WebSocket the client must stop playback and truncate the conversation item itself when an interruption fires.
Four different ways agents are billed
Agent billing is separate from text-to-speech billing almost everywhere, and the units are not comparable with each other any more than the latency figures are.
- Cartesia bills agents per minute in US dollars, and they explicitly do not draw on the credit balance used for text-to-speech. Two currencies, one account — as our pricing comparison records.
- Speechify bills agents per minute and text-to-speech per character, both drawing down one shared prepaid dollar balance.
- Azure meters its Voice Live API in tokens, tiered by the generative model chosen — not per call minute. The tier follows from the model rather than from a plan you select.
- Google bills input and output seconds, and documents a consequence worth knowing before you enable a feature: turning on barge-in increases billable duration, because input and output seconds are billed simultaneously during playback. It is the only vendor here that documents a capability making the bill go up.
Geography, and two traps in it
Where the compute runs matters more for a live conversation than any model figure, since network round-trip is a component that no model improvement removes. Six providers publish regional endpoints and a seventh names regions without publishing hostnames; the detail is where the traps are.
OpenAI’s regional endpoints do not move the processing for text-to-speech. Ten regional hostnames are published, but for the audio endpoint group containing text-to-speech, most regions offer regional storage only — only the United States and Europe show processing. Choosing a Tokyo endpoint for data residency does not put the synthesis compute in Tokyo.
Amazon Polly is not available in every region it has an endpoint for. Endpoints are published for 25 AWS regions, but each voice engine has its own, much narrower list — and generative voices, the newest engine, are restricted to 10 regions with an explicit closed-world clause. The endpoint list is not the availability list.
The rest, briefly:
- ElevenLabs publishes three named data-residency environments with exact hostnames, Enterprise-gated, and documents storage residency as separate from processing location.
- Murf publishes eleven regional hostnames plus a global router, and explicitly separates a residency-first endpoint choice from a latency-first one — the clearest framing of the trade-off here.
- Azure publishes a per-region hostname pattern and enforces region by credential scoping rather than by hostname alone.
- Google publishes three regional and multi-regional hostnames and attaches no performance claim to them, stating only that nothing else about the API changes. That restraint is a documentation virtue, not a gap.
- Cartesia names regions but publishes no regional hostnames, gates them to Enterprise by contact, and documents a relationship between region and latency without attaching a number or naming a metric.
- Speechify documents a single global base URL with no regional variant anywhere in its API reference.
- Resemble publishes global hostnames split by function rather than geography, and its API keys are explicitly not region-scoped. Its plan gating is on transport instead: the lowest-latency transport is restricted to Business plans and above.
How to measure this for yourself
Latency is the one property on this site you can measure yourself more easily than we can measure it for you, because the only measurement that matters is the one taken from where your users are, on your payload.
- Decide which metric you care about first. Almost certainly time-to-first-audio, measured from your application issuing the request to the first sample reaching a speaker. Vendors’ model-inference figures exclude the network, and the network is most of your budget.
- Measure from where your users are, not from your laptop and not from a cloud region next to the vendor. ElevenLabs’ own regional table moves by 50–100 ms across geographies.
- Use your real payload. Synthesis time scales with text length; a figure taken on a five-word string will not survive a paragraph.
- Report a distribution, not a best case. Run it a few hundred times and record p50 and p95. No vendor publishes a percentile, which is precisely why yours is worth having.
- Measure under your real concurrency. Queueing behaviour at your concurrency limit is a documented effect, and a single-request measurement will never show it.
- Use a persistent connection if you will in production. Murf documents that persistent connections avoid “30 - 80 ms of extra overhead per request” from repeated handshakes — connection strategy is worth as much as model choice.
- Record the model version and the date. These models change under fixed names, and no vendor dates its own speed claims.
Do that and you will have something no page on the internet currently has: a latency figure with a defined metric, a location, a payload, a load and a date.
Frequently asked questions
Which AI voice generator has the lowest latency?
We do not know, and the published numbers cannot answer it. They measure different quantities: model inference at one vendor, an undefined “latency” at another, time-to-first-audio at a third, and an ordinal rating at a fourth. Comparing them produces an ordering with no meaning.
Is 75 ms really faster than 90 ms?
Those two figures do not measure the same thing. The 75 ms is model inference latency, which its own vendor states excludes network round-trips and application overhead and is “always” smaller than what a user experiences. The 90 ms carries no metric definition at all. The comparison is not available.
What latency should I actually plan for?
Measure it yourself from your users’ location, on your payload, at your concurrency. The only published region-conditioned figures we found are ElevenLabs’, which give 100–150 ms TTFB for North America, Europe and South East Asia and 150–200 ms for South and North East Asia — both well above that vendor’s own 75 ms headline, and both vendor claims rather than our measurements.
Which providers can actually build a voice agent?
Seven document one: Amazon (through Lex V2), Azure, Cartesia, ElevenLabs, Google, OpenAI and Speechify. MiniMax and Resemble AI sell no agent product. Murf markets one with no developer documentation.
Does every agent support interruption?
Every documented agent product mentions it, but not all expose it as something you can configure. Amazon, Azure, ElevenLabs, Google, OpenAI and Speechify publish named parameters. Cartesia asserts it in prose without a configurable field in its agent configuration reference.
Why do two vendors publish no latency number at all?
Google attaches no figure to text-to-speech, and OpenAI publishes an ordinal “Speed” rating instead of a time. We treat neither as a weakness. A vendor that publishes no number is not asking you to rely on one.
Does choosing a nearby region make it faster?
Sometimes, and check what the region actually moves. OpenAI’s regional endpoints give storage residency for the audio group but processing only in the United States and Europe. Google explicitly states its regional endpoints change nothing but data location.
Will you ever rank this?
When we can measure it to a documented method with stored results you can check. Until then a ranking would be an ordering of marketing claims, and this page exists to explain why that is worthless.
Sources and what we could not verify
Every claim on this page was read from the vendor’s own documentation, marketing pages or API references on September 2, 2026. Every latency figure is reproduced as a vendor claim and none is endorsed, compared or converted. Where a vendor names a metric, we report the name; where it defines one, we quote the definition; where it does neither, we say so. Per-provider detail is in the ten reviews linked throughout. Integration limits, transports and concurrency are covered separately in our text-to-speech API comparison, and the method is on the methodology page.
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 2, 2026:
- Any latency figure on this page. We measured nothing. Every number is the vendor’s, published without a percentile, a date, a payload or a load, and we could not verify a single one.
- What most of these metrics mean. Cartesia, Speechify and most of Murf’s figures name a metric that no vendor page defines, or no metric at all.
- Cartesia’s current model. Sonic 3.6 is generally available and carries no latency figure; the widely quoted sub-90 ms belongs to Sonic 3.5.
- ElevenLabs’ TTFA. The vendor defines it as the number that matters and attaches no figure to it.
- Whether any vendor’s claim holds under load. Only Murf asserts a figure under concurrency, and an assertion is not a measurement.
- Murf’s agent product. Marketed, sold by sales contact, with no published developer documentation to assess.
- Whether regional endpoints reduce latency for any provider. Only Cartesia claims a relationship, without a number or a metric; Google explicitly claims none.
- How any of this sounds. No listening test has been run. Speed and quality are separate questions and this page answers neither.
Change log
September 2, 2026 — First publication. All claims read from official vendor pages on September 2, 2026 and dated accordingly. Published without a speed ranking, for the reason given at the top.
Published September 2, 2026 · This page reports no measurements of any kind. When our benchmark covers latency, it will publish a defined metric, a stated location, a stated payload and a distribution — and this page will be revisited with a dated entry (methodology). Corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.