Real-Time AI Voice Generators

✓ Docs-verified · Sep 2, 2026◌ Not audio-tested△ What we could not verify

Published September 2, 2026 · Every figure below was read from the vendor’s own pages on September 2, 2026. This site has run no latency measurement. Every number on this page is a vendor claim, labelled as one, and no table here is sorted by speed.

The short answer

We will not tell you which of these is fastest, and neither can anyone else from published figures. The numbers vendors publish measure different things, and most vendors never say which thing.

One vendor publishes three different latency numbers for the same model — and they need not conflict, because each names a different metric. ElevenLabs states ~75 ms, 100–150 ms and sub-500 ms for Flash v2.5 in three places.

That same vendor tells you the problem in its own documentation: “Latency in audio generation is deceptively simple-sounding but involves several distinct phenomena that are easy to conflate.” It then says its own headline figure is not the one that matters.

What is comparable is agent capability, and it is checkable. Does barge-in exist as a configurable parameter? Is turn detection parameterised? Is telephony documented? Seven providers answer yes, two sell no agent product at all, and one markets an agent with no developer documentation.

Basis: vendor documentation only, checked September 2, 2026 — no measurements, no accounts, no API calls (how we verify).

On this page

Why this page publishes no speed ranking

Ranking latency requires measuring it: the same input, the same transport, the same region, the same concurrency, repeated enough times to report a distribution rather than a best case. We have not done that, and until we do a ranking from this site would be an ordering of other people’s marketing.

It would also be wrong in a specific and demonstrable way. The published figures are not measurements of the same quantity. Putting ElevenLabs’ 75 ms next to Cartesia’s 90 ms produces an ordering, and the ordering is an artifact of which metric each vendor chose to publish rather than of how fast anything is. The next section shows that from the vendors’ own pages.

So this page does two things instead. It sets out what each vendor actually claims and what that claim measures, so you can stop treating those numbers as a league table. And it compares the things that are documented and checkable — interruption handling, turn detection, telephony and billing — which turn out to matter more for a working voice agent than any millisecond figure.

The metric problem

Below is every latency figure we found, with the metric the vendor named and whether the vendor defined it. The table is ordered alphabetically, not by number, and the numbers are not comparable with each other. That is the point of the table.

Published latency claims and the metrics behind them — vendor claims only, read September 2, 2026, ordered alphabetically
ProviderVendor’s figureMetric namedDefined by the vendor?
Cartesia“sub-90ms latency”None — the bare word “latency”No. A separate page instead describes streaming “the first byte of audio in about 90ms”
ElevenLabs“~75ms†”“model inference latency”Yes — and the vendor says it is not the number that matters
ElevenLabs100–150 ms / 150–200 ms by region“TTFB”No — named but never expanded or defined on that page
ElevenLabs“sub-500ms”“end-to-end latency”No — named, not defined on that page
Google CloudNo number published
Murf AI“55ms”“model inference” and “model latency”, used interchangeablyNo
Murf AI“~100ms” (docs) / “130ms” (marketing)“time-to-first-audio”Yes, on one page: “the delay between a synthesis request and the first audio frame”
OpenAINo number published“Speed” — an ordinal rating, “Fast” or “Medium”Not applicable — no time value is published at all
Speechify“<300ms”“TTFB” and “startup latency”, on different surfacesNo — neither term is expanded or defined anywhere we read

Read the second, third and fourth rows together. They are all ElevenLabs, and they are all about the same model family. 75 ms is model inference. 100–150 ms is TTFB for North America, Europe and South East Asia, rising to 150–200 ms for South and North East Asia. Sub-500 ms is end-to-end. Three numbers, three metrics, one product, and no contradiction between them — only one is the number a user would experience. We have verified none of the three; they are the vendor’s.

Cartesia is the opposite case. Its published figure carries no metric at all — “sub-90ms latency” — and its documentation landing page separately describes Sonic 3.5 as streaming “the first byte of audio in about 90ms”. Those are different quantities sharing a number. Meanwhile Sonic 3.6, which Cartesia announced as generally available, carries no latency figure at all: its page contains zero values of the form “nms”. The headline number belongs to the previous model.

Murf publishes the widest spread and the clearest definition. Its marketing pairs “55ms model inference” with “130ms end-to-end latency” on one page, its documentation says “~100ms time-to-first-audio” for Falcon 2, and one Murf page defines time-to-first-audio properly as the delay between a synthesis request and the first audio frame. It is the only vendor here besides ElevenLabs to define anything.

And Murf defines the same metric two different ways. Its latency documentation names the metric “Time to First Audio Byte” and operationalises it against a concrete curl field, time_starttransfer. Its benchmark page defines time-to-first-audio as the delay between a synthesis request and the first audio frame. A byte and a frame are not the same event, so two Murf pages using the same acronym are measuring two different things. That is the metric problem happening inside a single vendor.

Three vendors publish no number at all. Google Cloud attaches no latency figure to text-to-speech anywhere we read, and neither does MiniMax on either of its platforms. OpenAI publishes an ordinal instead — a “Speed” label reading “Fast” or “Medium”. Neither absence is a weakness; a vendor that publishes no number is not making a claim you have to discount.

The vendor that explains the problem itself

The most useful document in this entire subject is published by ElevenLabs, and it undercuts ElevenLabs’ own marketing. It is worth quoting at length because no summary improves on it.

The page opens: “Latency in audio generation is deceptively simple-sounding but involves several distinct phenomena that are easy to conflate.”

It then defines the two quantities that get conflated. First: “Model inference latency is the time the model spends generating audio. ElevenLabs Flash models achieve ~75ms model inference for typical short inputs. This is an internal measurement, excluding network round-trips and application overhead.”

And second: “Time-to-first-audio (TTFA) is the elapsed time from when your application initiates a request to when the first audio sample actually plays for the end user. This is almost always the number that matters for user experience, and it is always larger — often substantially larger — than model inference latency alone.”

Elsewhere the vendor is equally direct about its headline figure: “75ms refers to model inference time only. Actual end-to-end latency will vary with factors such as your location & endpoint type used.” And of that same figure: “It is a useful reference point for comparing models, not a guarantee for every request.”

So the vendor that publishes the lowest headline number in this market also publishes, in its own documentation, the statement that the number is not what a user experiences and that the figure which is — TTFA — is always larger. No TTFA figure is attached to it anywhere on that page.

The dagger on “~75ms†” resolves to a scope note rather than a definition: “† Excluding application & network latency”. Two of the three places its ~280 ms figure for the conversational model appears carry no dagger at all, so the exclusion travels with the number only some of the time.

None of this makes ElevenLabs’ claims dishonest. It makes them precise in a way the rest of the market’s are not, and it means the one vendor documenting the problem is the one whose headline number gets quoted most often without it.

What no vendor publishes

Across every page we read, for all ten providers, we found none of the following attached to any latency figure:

A figure with no percentile, no date, no payload, no load and no location is not something you can plan a system around. It is a marketing number, and treating it as anything else is the mistake this page exists to prevent.

What is comparable: agent capability

Here is the constructive half. Whether a provider documents interruption handling as a configurable parameter is not a matter of opinion or measurement — it either appears in the API reference or it does not. The same goes for turn detection and telephony. These are the things that decide whether a voice agent works, and unlike latency they are checkable today.

Documented voice-agent capability — read September 2, 2026, ordered alphabetically
ProviderAgent productInterruptionTurn detectionTelephony
AmazonAmazon Lex V2, which speaks using Polly voicesallowInterrupt, on by default, at prompt and per-attempt levelTwo required millisecond parametersDocumented
Azure SpeechVoice Live APIinterrupt_response, plus a published filler-word filter to reduce false triggersFour named types with published defaults and rangesDocumented
CartesiaManaged Agents / LineAsserted in prose; not a configurable field in the agent configuration referenceParameterised on the standalone speech-to-text endpoint, not on the agent productOwn SIP endpoint, bring-your-own-carrier
ElevenLabsElevenAgents“Interruption handling”, configurable on/off“Take turn after silence” plus “Turn eagerness” with three modesSIP trunking, inbound and outbound
Google CloudConversational Agents (separate billed product)“Barge-in”, a per-level toggleEnd-of-speech sensitivity, smart endpointing, no-speech timeoutHosts phone numbers directly
MiniMaxNone
Murf AIMarketed, no developer documentationMarketing prose only — no term of art, no parameterDescribed as customisable; no parameters publishedNot documented
OpenAIRealtime API, plus RealtimeAgent in the Agents SDK“barge-in”, via turn_detection.interrupt_responseserver_vad and semantic_vad — silence-based and classifier-basedOwn SIP endpoint; does not sell phone numbers
Resemble AINone
SpeechifySpeechifyAI Agents (beta)interruption_sensitivity, a three-level enum with published thresholdsturn_handling object with published default wait windowsPhone numbers via LiveKit, Twilio or bring-your-own SIP

The gap between marketing and documentation is where this table earns its place. Murf markets AI voice agents; its API documentation index contains no agent endpoints and no conversational API. Cartesia’s documentation says its agent “handles turn-taking and interruptions” in prose, but interruption is not exposed as a configurable field in the agent configuration reference, and its rich turn-detection parameters belong to a different product. Both may work well. Neither publishes what you would need to tune it.

Two vendors sell no agent product at all. MiniMax documents text-to-speech, async long-form, cloning, voice design and voice management — and no realtime speech-to-speech surface. The word “Agent” appears throughout its documentation but never in a voice sense. Resemble AI likewise, which is unsurprising given it no longer sells voice generation.

The most detailed interruption controls belong to the platform vendors and OpenAI. Azure publishes an English filler-word list whose stated purpose is reducing false barge-in triggers — a level of specificity nobody else reaches. OpenAI documents two turn-detection modes, one silence-based and one classifier-based, and is explicit that on WebSocket the client must stop playback and truncate the conversation item itself when an interruption fires.

Four different ways agents are billed

Agent billing is separate from text-to-speech billing almost everywhere, and the units are not comparable with each other any more than the latency figures are.

Geography, and two traps in it

Where the compute runs matters more for a live conversation than any model figure, since network round-trip is a component that no model improvement removes. Six providers publish regional endpoints and a seventh names regions without publishing hostnames; the detail is where the traps are.

OpenAI’s regional endpoints do not move the processing for text-to-speech. Ten regional hostnames are published, but for the audio endpoint group containing text-to-speech, most regions offer regional storage only — only the United States and Europe show processing. Choosing a Tokyo endpoint for data residency does not put the synthesis compute in Tokyo.

Amazon Polly is not available in every region it has an endpoint for. Endpoints are published for 25 AWS regions, but each voice engine has its own, much narrower list — and generative voices, the newest engine, are restricted to 10 regions with an explicit closed-world clause. The endpoint list is not the availability list.

The rest, briefly:

How to measure this for yourself

Latency is the one property on this site you can measure yourself more easily than we can measure it for you, because the only measurement that matters is the one taken from where your users are, on your payload.

  1. Decide which metric you care about first. Almost certainly time-to-first-audio, measured from your application issuing the request to the first sample reaching a speaker. Vendors’ model-inference figures exclude the network, and the network is most of your budget.
  2. Measure from where your users are, not from your laptop and not from a cloud region next to the vendor. ElevenLabs’ own regional table moves by 50–100 ms across geographies.
  3. Use your real payload. Synthesis time scales with text length; a figure taken on a five-word string will not survive a paragraph.
  4. Report a distribution, not a best case. Run it a few hundred times and record p50 and p95. No vendor publishes a percentile, which is precisely why yours is worth having.
  5. Measure under your real concurrency. Queueing behaviour at your concurrency limit is a documented effect, and a single-request measurement will never show it.
  6. Use a persistent connection if you will in production. Murf documents that persistent connections avoid “30 - 80 ms of extra overhead per request” from repeated handshakes — connection strategy is worth as much as model choice.
  7. Record the model version and the date. These models change under fixed names, and no vendor dates its own speed claims.

Do that and you will have something no page on the internet currently has: a latency figure with a defined metric, a location, a payload, a load and a date.

Frequently asked questions

Which AI voice generator has the lowest latency?

We do not know, and the published numbers cannot answer it. They measure different quantities: model inference at one vendor, an undefined “latency” at another, time-to-first-audio at a third, and an ordinal rating at a fourth. Comparing them produces an ordering with no meaning.

Is 75 ms really faster than 90 ms?

Those two figures do not measure the same thing. The 75 ms is model inference latency, which its own vendor states excludes network round-trips and application overhead and is “always” smaller than what a user experiences. The 90 ms carries no metric definition at all. The comparison is not available.

What latency should I actually plan for?

Measure it yourself from your users’ location, on your payload, at your concurrency. The only published region-conditioned figures we found are ElevenLabs’, which give 100–150 ms TTFB for North America, Europe and South East Asia and 150–200 ms for South and North East Asia — both well above that vendor’s own 75 ms headline, and both vendor claims rather than our measurements.

Which providers can actually build a voice agent?

Seven document one: Amazon (through Lex V2), Azure, Cartesia, ElevenLabs, Google, OpenAI and Speechify. MiniMax and Resemble AI sell no agent product. Murf markets one with no developer documentation.

Does every agent support interruption?

Every documented agent product mentions it, but not all expose it as something you can configure. Amazon, Azure, ElevenLabs, Google, OpenAI and Speechify publish named parameters. Cartesia asserts it in prose without a configurable field in its agent configuration reference.

Why do two vendors publish no latency number at all?

Google attaches no figure to text-to-speech, and OpenAI publishes an ordinal “Speed” rating instead of a time. We treat neither as a weakness. A vendor that publishes no number is not asking you to rely on one.

Does choosing a nearby region make it faster?

Sometimes, and check what the region actually moves. OpenAI’s regional endpoints give storage residency for the audio group but processing only in the United States and Europe. Google explicitly states its regional endpoints change nothing but data location.

Will you ever rank this?

When we can measure it to a documented method with stored results you can check. Until then a ranking would be an ordering of marketing claims, and this page exists to explain why that is worthless.

Sources and what we could not verify

Every claim on this page was read from the vendor’s own documentation, marketing pages or API references on September 2, 2026. Every latency figure is reproduced as a vendor claim and none is endorsed, compared or converted. Where a vendor names a metric, we report the name; where it defines one, we quote the definition; where it does neither, we say so. Per-provider detail is in the ten reviews linked throughout. Integration limits, transports and concurrency are covered separately in our text-to-speech API comparison, and the method is on the methodology page.

What we could not verify

Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 2, 2026:

Change log

September 2, 2026 — First publication. All claims read from official vendor pages on September 2, 2026 and dated accordingly. Published without a speed ranking, for the reason given at the top.

Published September 2, 2026 · This page reports no measurements of any kind. When our benchmark covers latency, it will publish a defined metric, a stated location, a stated payload and a distribution — and this page will be revisited with a dated entry (methodology). Corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.