Methodology: how we verify what we publish
Two rules govern every page on this site. First: every fact that can change — a price, a model name, a free-tier limit, a licensing term — is checked against the vendor's own current documentation on the day we write, and carries the date we checked it. Second: our standardized audio benchmark exists today as a finished design, not as a running system. Until it runs, no page on this site claims a listening test, and no page carries a numeric score.
Updated August 2026 · Audio benchmark status: designed, not yet running. No benchmark audio exists, and no numeric scores appear anywhere on this site.
How facts are verified today
We do not write from memory, and we do not copy claims from other roundups. Every changing fact is looked up live, at writing time, from sources in this order:
| Priority | Source | When it is used |
|---|---|---|
| 1 — Official | The provider's own documentation, pricing pages, API references, and terms of service | Always first. The default for every claim. |
| 2 — Primary technical | Official changelogs, SDK repositories, model cards, status pages, published research | When official pages do not answer, or to corroborate them. |
| 3 — Secondary | Credible third-party reviews and comparisons | Only when the first two do not answer — and never for licensing terms. |
Every changing fact carries its date
A price, quota, or feature claim is only as good as the day it was checked. Each one appears on the page with its verification date next to it — in the same sentence, table row, or labeled metadata block — so you always know how fresh the claim is.
Licensing is quoted, not paraphrased
Commercial-use and output-ownership statements come only from the provider's current official terms, and we quote the operative language verbatim wherever practical. Rights verified for one plan are never extended to another: a paid plan's commercial rights say nothing about the free tier. Where the terms are silent or ambiguous, the page says so instead of guessing.
Disagreements are stated, not hidden
When sources conflict, the more authoritative and more recently dated official source wins. If the conflict survives that test, the page states the range or the uncertainty openly rather than silently picking one number.
Vendor claims are labeled as vendor claims
Nothing a vendor says about its own product is presented as something we established ourselves. A vendor's claim is always attributed as the vendor's claim. The full set of writing rules is in our editorial policy.
Where audio testing stands today
We have designed a standardized audio benchmark — the same short prompts, generated under the same conditions by every provider, so results can be compared directly. That system is not yet running. No benchmark audio has been generated, and we have not tested any product first-hand.
Until the pipeline runs, three things hold everywhere on this site:
- No page claims that we tested, tried, or listened to anything. Any such wording would be fabricated, and we treat it as publish-blocking.
- Every claim is verified against official vendor documentation, with dated sources, as described above.
- No numeric scores are published. Score bands and weights are finalized only against real measurements; printing precise-looking numbers before those measurements exist would be false precision. Verdicts until then are qualitative and documentation-based, and pages say so.
When the benchmark starts running, this page will be updated to record it — including the finalized scoring rubric and the frozen prompt texts.
The audio benchmark, as designed
The point of the benchmark is comparability. Every provider will generate the same short texts, in the same language, under the same conditions — so any difference a listener hears will come from the products, not from the tests.
Three standard prompts
| Test | What it measures |
|---|---|
| Naturalness | Pacing, phrasing, pauses, clarity, and overall realism — whether the voice sounds like effortless speech. |
| Precision | Hard-to-say content: numbers, dates and times, currency amounts, and mixed letter-and-number codes. |
| Expression (optional) | Energy, emphasis, and emotional control, where a model supports it. |
The exact spoken texts are frozen when the pipeline activates, after their generated durations have been calibrated across several providers. Any later change to a frozen prompt is versioned and documented on this page, and samples from different prompt versions are never compared side by side as if identical.
A small, fixed sample budget
| Rule | Design value |
|---|---|
| Samples per provider | Two required — Naturalness and Precision — plus the optional Expression sample. |
| Length per sample | Roughly 10–20 seconds; never intentionally long. |
| Total per provider | Capped at 45 seconds of core audio. |
| Generation frequency | Once per provider and model version, unless a retest is justified by a material change. |
The budget is small on purpose: short, identical samples are cheap to standardize, quick to compare, and leave nowhere for weaknesses to hide.
Same conditions for every provider
- The same spoken text across all providers in a run.
- The same language and locale — English (en-US) by default.
- A neutral adult voice with similar characteristics across providers where possible; where exact matching is impossible, the mismatch is documented with the sample.
- Default settings, unless a change is needed to reach normal-quality output — and every deviation is recorded. One provider is never tuned while its rivals sit at defaults.
- Provider-supplied voices only. We will never clone a real person's voice without explicit rights and consent.
Generate once, reuse everywhere
Each provider's samples will be generated once per model version and stored with full metadata: model, voice, prompt version, duration, generation date, and what the audio actually cost to produce. Every page that presents that provider's benchmark reuses the same stored samples — nothing is regenerated per article. Samples are regenerated only when something material changes: the model or voice, the pricing, or the methodology itself. Older versions are kept for transparency.
Scoring: nothing numeric yet
No scores exist on this site today, for the reason above: a scoring methodology is only finalized against real measurements. When scoring goes live, these rules bind it:
- Score bands are defined before testing begins, not fitted afterwards.
- A published score never carries more precision than the method can support.
- Raw measurements are stored behind every published number.
- Editorial judgment and objective measurement are kept separate and labeled.
- Methodology changes are documented here, and a historic score is never changed silently.
How pages stay current
Facts about AI voice generators change quickly, so pages are rechecked on a schedule matched to how fast their facts move — pricing most often, slower-moving guides less often. Three habits keep the dates honest:
- Dated facts. Every volatile fact keeps its verification date beside it, as described above.
- Truthful "updated" dates. A displayed updated date — including the dates search engines see — changes only when something meaningful changed on the page. Never to look fresh.
- Update notes. When a meaningful change ships — a price move, a model release, a methodology revision — the affected page records a short dated note saying what changed.
If we got a fact wrong, that is handled — visibly — under our corrections policy. And no payment or vendor relationship influences what we publish; how the site is funded is set out in how we make money.