Methodology: how we verify what we publish

Two rules govern every page on this site. First: every fact that can change — a price, a model name, a free-tier limit, a licensing term — is checked against the vendor's own current documentation on the day we write, and carries the date we checked it. Second: our standardized audio benchmark exists today as a finished design, not as a running system. Until it runs, no page on this site claims a listening test, and no page carries a numeric score.

Updated August 2026 · Audio benchmark status: designed, not yet running. No benchmark audio exists, and no numeric scores appear anywhere on this site.

How facts are verified today

We do not write from memory, and we do not copy claims from other roundups. Every changing fact is looked up live, at writing time, from sources in this order:

Source order for every changing fact
PrioritySourceWhen it is used
1 — Official The provider's own documentation, pricing pages, API references, and terms of service Always first. The default for every claim.
2 — Primary technical Official changelogs, SDK repositories, model cards, status pages, published research When official pages do not answer, or to corroborate them.
3 — Secondary Credible third-party reviews and comparisons Only when the first two do not answer — and never for licensing terms.

Every changing fact carries its date

A price, quota, or feature claim is only as good as the day it was checked. Each one appears on the page with its verification date next to it — in the same sentence, table row, or labeled metadata block — so you always know how fresh the claim is.

Licensing is quoted, not paraphrased

Commercial-use and output-ownership statements come only from the provider's current official terms, and we quote the operative language verbatim wherever practical. Rights verified for one plan are never extended to another: a paid plan's commercial rights say nothing about the free tier. Where the terms are silent or ambiguous, the page says so instead of guessing.

Disagreements are stated, not hidden

When sources conflict, the more authoritative and more recently dated official source wins. If the conflict survives that test, the page states the range or the uncertainty openly rather than silently picking one number.

Vendor claims are labeled as vendor claims

Nothing a vendor says about its own product is presented as something we established ourselves. A vendor's claim is always attributed as the vendor's claim. The full set of writing rules is in our editorial policy.

Where audio testing stands today

We have designed a standardized audio benchmark — the same short prompts, generated under the same conditions by every provider, so results can be compared directly. That system is not yet running. No benchmark audio has been generated, and we have not tested any product first-hand.

Until the pipeline runs, three things hold everywhere on this site:

When the benchmark starts running, this page will be updated to record it — including the finalized scoring rubric and the frozen prompt texts.

The audio benchmark, as designed

The point of the benchmark is comparability. Every provider will generate the same short texts, in the same language, under the same conditions — so any difference a listener hears will come from the products, not from the tests.

Three standard prompts

The standard prompt suite
TestWhat it measures
Naturalness Pacing, phrasing, pauses, clarity, and overall realism — whether the voice sounds like effortless speech.
Precision Hard-to-say content: numbers, dates and times, currency amounts, and mixed letter-and-number codes.
Expression (optional) Energy, emphasis, and emotional control, where a model supports it.

The exact spoken texts are frozen when the pipeline activates, after their generated durations have been calibrated across several providers. Any later change to a frozen prompt is versioned and documented on this page, and samples from different prompt versions are never compared side by side as if identical.

A small, fixed sample budget

Per-provider audio budget
RuleDesign value
Samples per provider Two required — Naturalness and Precision — plus the optional Expression sample.
Length per sample Roughly 10–20 seconds; never intentionally long.
Total per provider Capped at 45 seconds of core audio.
Generation frequency Once per provider and model version, unless a retest is justified by a material change.

The budget is small on purpose: short, identical samples are cheap to standardize, quick to compare, and leave nowhere for weaknesses to hide.

Same conditions for every provider

Generate once, reuse everywhere

Each provider's samples will be generated once per model version and stored with full metadata: model, voice, prompt version, duration, generation date, and what the audio actually cost to produce. Every page that presents that provider's benchmark reuses the same stored samples — nothing is regenerated per article. Samples are regenerated only when something material changes: the model or voice, the pricing, or the methodology itself. Older versions are kept for transparency.

Scoring: nothing numeric yet

No scores exist on this site today, for the reason above: a scoring methodology is only finalized against real measurements. When scoring goes live, these rules bind it:

How pages stay current

Facts about AI voice generators change quickly, so pages are rechecked on a schedule matched to how fast their facts move — pricing most often, slower-moving guides less often. Three habits keep the dates honest:

If we got a fact wrong, that is handled — visibly — under our corrections policy. And no payment or vendor relationship influences what we publish; how the site is funded is set out in how we make money.