Best AI Voice Generators in 2026: 10 Providers Compared

Updated August 2026 · Every fact verified against official vendor documentation on August 11, 2026.

The fast answer. For expressive creator voices with self-serve cloning, ElevenLabs leads — its commercial license starts on the $6/month Starter plan (elevenlabs.io/pricing, checked Aug 11, 2026). For developers, Google Cloud Text-to-Speech is the verified price floor at $0.004 per 1,000 characters (official pricing, checked Aug 11, 2026). On a tight budget, Cartesia's $5/month Pro plan includes a commercial license and instant voice cloning (cartesia.ai/pricing, checked Aug 11, 2026).

Decide faster: Compare pricing · Find your fit · How we verify

Every price, licensing term, and feature claim on this page was read from official vendor documentation on August 11, 2026, with the source linked beside the fact. Our standardized audio benchmark is designed but not yet running — this page contains no listening-test claims and no scores.

On this page

The ranking at a glance

The order below is an editorial judgment built from documented pricing, licensing clarity, and feature coverage for these text to speech (TTS) products — not from listening tests, which we have not run yet. How we chose explains the basis; each provider name links to its full assessment.

10 AI voice generators, ranked by documented fit. Every fact verified against the linked official sources on August 11, 2026; no provider has been audio-tested by us.
# Provider Best for Price signal Free-tier commercial use Verification status
1 ElevenLabs Expressive creator voices and self-serve voice cloning Plans from $6/mo; API $0.05–$0.10/1k chars (pricing, API pricing, checked Aug 11, 2026) No — free plan is non-commercial (ToS §1(c)), attribution required (billing docs, both checked Aug 11, 2026) Docs-verified Aug 11, 2026
2 Google Cloud Text-to-Speech Developers at volume; lowest verified list price From $0.004/1k chars, pay-as-you-go only (pricing, checked Aug 11, 2026) Terms draw no free-vs-paid distinction (cloud.google.com/terms, checked Aug 11, 2026) Docs-verified Aug 11, 2026
3 Murf AI Studio editing workflow with quotable commercial rights Studio $29–$199/mo; API $0.03/1k chars (pricing, API plans, checked Aug 11, 2026) No usable free output — trial is self-evaluation only, no downloads (help.murf.ai, ToS §10.7, checked Aug 11, 2026) Docs-verified Aug 11, 2026
4 Speechify Cheapest entry-level character-billed API subscription ($10/mo) API from $10/mo for 1M chars (speechify.ai/pricing, checked Aug 11, 2026) Studio free: “No commercial usage rights”; API terms silent (pricing-studio, API terms, checked Aug 11, 2026) Docs-verified Aug 11, 2026
5 Amazon Polly AWS developers; cleanest output-ownership terms $0.004–$0.100/1k chars by engine, pay-as-you-go (pricing, checked Aug 11, 2026) No tier-scoped restriction found (FAQ, service terms, checked Aug 11, 2026) Docs-verified Aug 11, 2026
6 Azure Speech in Foundry Tools Enterprise scale, compliance, and the largest claimed catalog among the cloud platforms here $0.015/1k chars pay-as-you-go, to $0.006/1k on commitments (Retail Prices API, checked Aug 11, 2026) No — the commercial grant covers the paid tier only (Product Terms, checked Aug 11, 2026) Docs-verified Aug 11, 2026
7 OpenAI Instruction-steerable speech inside the OpenAI stack tts-1 $0.015/1k chars; current flagship model token-billed (pricing, checked Aug 11, 2026) No free TTS tier exists (pricing, checked Aug 11, 2026) Docs-verified Aug 11, 2026
8 Cartesia Budget real-time agents; cheapest commercial entry point Pro $5/mo incl. commercial license and cloning (pricing, checked Aug 11, 2026) No — free tier is expressly non-commercial (ToS §4.1, checked Aug 11, 2026) Docs-verified Aug 11, 2026
9 MiniMax Cheap API long-form and the lowest documented cloning fee $0.06–$0.10/1k chars pay-as-you-go (pricing, checked Aug 11, 2026) No free API tier documented; consumer app is non-commercial (consumer ToS, checked Aug 11, 2026) Docs-verified Aug 11, 2026
10 Resemble AI Watermarked, consent-tooled enterprise voice; self-hosting via open source TTS list pricing not published (pricing page, checked Aug 11, 2026) Terms silent on free-tier rights (ToS, checked Aug 11, 2026) Docs-verified Aug 11, 2026

Not sure which fits you? Jump to who should choose what for recommendations by use case, see how the numbers were built in pricing, normalized, or read how we chose first.

How we chose — and what we do not claim

Every changing fact on this page — every price, model name, quota, free-tier limit, and licensing term for these TTS products — was read from the vendor's own current documentation on August 11, 2026, and is cited in the same sentence or table cell where it appears. Licensing language is quoted from the operative terms wherever practical, and rights verified for one plan are never extended to another. The full protocol is on our methodology page; the writing rules are in our editorial policy.

Just as important is what this page does not claim. We have not generated audio from any of these products, created accounts with them, or run listening tests. Our standardized audio benchmark — the same short prompts generated under the same conditions by every provider — exists as a finished design, not a running system (methodology explains it). Until it runs, this page carries no naturalness verdicts, no "we tested" claims, and no scores of any kind. Where a vendor makes a quality or latency claim, it is labeled as the vendor's claim.

The ranking order is therefore an editorial judgment about documented fit: how many kinds of buyer each product serves well on verified pricing, licensing clarity, and feature coverage — with ties broken toward the provider whose terms answer the questions a buyer must ask. A provider whose terms are silent on commercial use, or whose pricing cannot be found on any public page, ranks below one that documents both, regardless of how its voices might sound. When the audio benchmark runs, this page will be refreshed with measured comparisons, and the change log will say so.

Where official sources disagree with each other — and several do — the disagreement is stated openly rather than silently averaged; see what we could not verify. One familiar name is absent from the list entirely: PlayHT. The reason is in the FAQ.

Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order. How the site is funded is set out in how we make money; factual errors are handled under our corrections policy.

Pricing, normalized — and how to read it

AI voice generator pricing splits into two families that resist direct comparison. Creator-facing tools sell subscriptions denominated in credits, "voice generation time", or points; API vendors bill per character. Only some vendors publish a conversion between their unit and characters — where one exists, the table normalizes to USD per 1,000 characters; where the vendor bills by time with no character conversion, the row is normalized per minute and says so; where no conversion exists at all, the cell says "not normalizable" rather than inventing one.

The spread is enormous: verified list prices run from $0.004 per 1,000 characters (Google Standard/WaveNet; Amazon Polly Standard, both checked Aug 11, 2026) to $0.160 per 1,000 characters (Google Studio voices, checked Aug 11, 2026) — a 40× range before creator subscriptions, whose plan-effective rates run higher still. Three reading rules:

AI voice generator pricing, normalized to list-price USD per 1,000 characters where the vendor bills by characters or publishes a character conversion — otherwise per minute, stated in the cell. All prices verified against official vendor pages on August 11, 2026.
Provider (pricing model) Billing basis Cheapest paid tier Normalized list price (formula shown) Free tier — quota and commercial use
ElevenLabs — subscription plans Credits; vendor states 1 credit = 1 character on Multilingual v2 models (elevenlabs.io/pricing, checked Aug 11, 2026) Starter, $6/mo, 30,000 credits, includes commercial license (elevenlabs.io/pricing, checked Aug 11, 2026) Starter: $6 ÷ 30,000 credits × 1,000 = $0.200/1k chars; Creator: $22 ÷ 121,000 × 1,000 = $0.182; Pro: $99 ÷ 600,000 × 1,000 = $0.165 (Multilingual-v2 basis; Flash costs 0.5–1 credit/char, exact rate unpublished) (elevenlabs.io/pricing, checked Aug 11, 2026) 10,000 credits/mo, $0. Non-commercial only, attribution required (ToS §1(c), updated Mar 31, 2026; billing docs, both checked Aug 11, 2026)
ElevenLabs — API list price USD per 1,000 characters, by model (elevenlabs.io/pricing/api, checked Aug 11, 2026) Pay-as-you-go at list rates, no commitment (elevenlabs.io/pricing/api, checked Aug 11, 2026) Multilingual v2 / v3: $0.10/1k chars; Flash / Turbo: $0.05/1k chars (vendor list price, no computation needed). Note: this page states plan inclusions in characters that are roughly 1.65×–2× (depending on plan) the credit quotas on the main pricing page — the two official pages disagree and must be attributed separately (elevenlabs.io/pricing/api, checked Aug 11, 2026) API free column: 10,000 Multilingual / 20,000 Flash chars (elevenlabs.io/pricing/api, checked Aug 11, 2026); same non-commercial terms as the free plan
Google Cloud Text-to-Speech Characters (classic tiers; SSML tags count, multibyte = 1 char); tokens for Gemini-TTS (cloud.google.com pricing, checked Aug 11, 2026) Pay-as-you-go only, no plans; billing account required even for free usage (cloud.google.com pricing, checked Aug 11, 2026) Standard / WaveNet: $4 / 1M ÷ 1,000 = $0.004/1k chars (cheapest verified list price in this table); Neural2: $16 / 1M = $0.016/1k; Chirp 3 HD: $30 / 1M = $0.030/1k; Instant Custom Voice: $60 / 1M = $0.060; Studio: $160 / 1M = $0.160. Gemini-TTS is token-billed; vendor publishes 25 audio tokens/sec, so output normalizes per minute only: $10/1M × 1,500 tokens = $0.015/min (2.5 Flash); $20/1M × 1,500 = $0.030/min (2.5 Pro, 3.1 Flash preview), plus unnormalizable text-input tokens (cloud.google.com pricing, checked Aug 11, 2026) Always-free monthly bands: 4M chars (Standard/WaveNet), 1M (Chirp 3 HD, Studio, Neural2, Polyglot); none for Gemini-TTS/ICV; overage auto-bills; terms silent on any free-vs-paid commercial distinction; $300/90-day trial requires a payment method (pricing; free tier docs, both checked Aug 11, 2026)
Murf AI — Studio plans Voice Generation Time (VGT), sold in hours; no vendor characters conversion published — normalized per minute (help.murf.ai VGT article, checked Aug 11, 2026) Creator Lite, $29/mo, 2 h VGT/mo ($23/mo equivalent on the $276 annual plan) (Murf's production pricing endpoint murf.ai/pricing data, checked Aug 11, 2026) Per minute: Creator Lite $29 ÷ 120 min = $0.242/min (annual: $276 ÷ 1,440 = $0.192/min); Business Plus $199 ÷ 1,200 = $0.166/min (annual $1,908 ÷ 14,400 = $0.133/min). Vendor's own approximation: ~6 min VGT per 1,000 words (help.murf.ai, checked Aug 11, 2026) Free trial: 10 min VGT, no card required, no downloads; ToS limits trials to “self-evaluation” (§10.7) (help.murf.ai; ToS, Dec 19, 2024, both checked Aug 11, 2026)
Murf AI — API pay-as-you-go Characters, sold in 1,000-character blocks (help.murf.ai API plans, checked Aug 11, 2026) Pay-as-you-go, $2 minimum purchase (help.murf.ai, checked Aug 11, 2026) $0.30 per 10,000-char unit ÷ 10 = $0.03/1k chars (help.murf.ai, checked Aug 11, 2026) API free trial: 100,000 characters, 1 API key, $0; card requirement not stated (help.murf.ai, checked Aug 11, 2026)
Speechify — API (SpeechifyAI) Characters for TTS; minutes for voice agents (speechify.ai/pricing, checked Aug 11, 2026) Starter, $10/mo, 1M TTS characters included (speechify.ai/pricing, checked Aug 11, 2026) Starter: $10 ÷ 1,000,000 × 1,000 = $0.010/1k chars (overage $10/1M = $0.010); Pro overage $8/1M = $0.008/1k; Scale overage $6/1M = $0.006/1k — the cheapest entry-level character-billed subscription in this table ($10/mo), with Scale overage tying Azure's largest commitment tier at $0.006/1k (speechify.ai/pricing, checked Aug 11, 2026) 50K chars/mo hard cap, $0; supplemental API terms are silent on free-vs-paid commercial rights (speechify.ai/pricing; API supplemental terms, Jun 10, 2024, both checked Aug 11, 2026)
Speechify — Studio Studio credits; vendor conversion 1 credit = 1 second of voiceover; no characters conversion — normalized per minute (speechify.com/pricing-studio, checked Aug 11, 2026) Studio Starter, $100/year, 86,400 credits (= 24 h voiceover/yr); monthly price not displayed (speechify.com/pricing-studio, checked Aug 11, 2026) Per minute: Starter $100 ÷ 1,440 min = $0.069/min; Creator $300 ÷ 5,760 min = $0.052/min (speechify.com/pricing-studio, checked Aug 11, 2026) 600 credits (= 10 min voiceover), $0; page states “No commercial usage rights” and no voice cloning (speechify.com/pricing-studio, checked Aug 11, 2026)
Speechify — app Premium (caution row) Words: 150,000 words/mo contractual floor; 1,000,000 words/mo guaranteed through 2026; no characters conversion — per 1,000 words only (speechify.com/usage-limits, checked Aug 11, 2026) Premium, $29/mo (annual figure not rendered — we do not print one) (speechify.com/pricing, checked Aug 11, 2026) $29 ÷ 150 (thousand words) = $0.193/1k words at the contractual floor. NOT comparable to commercial-use pricing: Premium is expressly non-commercial (ToS §4.4) (speechify.com/terms, Jul 1, 2025, checked Aug 11, 2026) App free: 10 “robotic sounding voices”, up to 1.5× speed, $0; non-commercial per ToS §4.4 (speechify.com/pricing, checked Aug 11, 2026)
Amazon Polly Characters, pay-as-you-go only, no plans (aws.amazon.com/polly/pricing, checked Aug 11, 2026) Pay-as-you-go at list rates; no subscription floor; AWS account with valid payment method required (AWS free-tier FAQ, checked Aug 11, 2026) Standard: $4.00 / 1M ÷ 1,000 = $0.004/1k chars; Neural: $16 / 1M = $0.016/1k; Generative: $30 / 1M = $0.030/1k; Long-Form: $100 / 1M = $0.100/1k. Vendor equivalence 1M chars ≈ 23 h 8 min supports a per-minute view (Neural ≈ $0.0115/min) (aws.amazon.com/polly/pricing, checked Aug 11, 2026) Per-engine monthly free tier (Standard 5M; Neural 1M; Long-Form 500K; Generative 100K chars — the latter three “for the first 12 months”; Standard's duration is stated inconsistently across AWS pages); accounts opened on/after Jul 15, 2025 get a $200-credit regime instead/alongside (precedence unstated). No tier-scoped commercial restriction found (pricing; FAQ, both checked Aug 11, 2026)
Azure Speech in Foundry Tools Characters per 1M (spaces/punctuation count; CJK characters bill as two) (Learn TTS overview, checked Aug 11, 2026) Pay-as-you-go S0, no subscription floor; Azure account requires a non-prepaid card for identity verification (azure.com account page, checked Aug 11, 2026) Standard neural: $15 / 1M ÷ 1,000 = $0.015/1k chars; Neural HD: $22 / 1M = $0.022/1k; commitment tiers: $960 ÷ 80,000 (k chars) = $0.012/1k down to $24,000 ÷ 4,000,000 = $0.006/1k. Dollar figures are from Microsoft's official Retail Prices API (eastus) because the pricing page rendered “$-” placeholders; Learn docs corroborate the $15/1M rate (prices.azure.com; Learn quotas doc, both checked Aug 11, 2026) F0: 0.5M neural characters/mo, $0. The commercial-use grant for prebuilt neural voice output covers “the paid tier TTS Service only” — F0 output is outside it (Microsoft Product Terms, checked Aug 11, 2026)
OpenAI — Audio API Characters for tts-1 / tts-1-hd; tokens for gpt-4o-mini-tts (developers.openai.com pricing, checked Aug 11, 2026) Pay-as-you-go only — prepaid credits, no subscription; card effectively required before any use (OpenAI help: prepaid billing, checked Aug 11, 2026) tts-1: $15 / 1M chars ÷ 1,000 = $0.015/1k chars; tts-1-hd: $30 / 1M ÷ 1,000 = $0.030/1k chars. gpt-4o-mini-tts is token-billed ($0.60/1M text-in, $12/1M audio-out) with no vendor token→character or token→minute conversion published — not normalized (developers.openai.com pricing, checked Aug 11, 2026) No free TTS tier exists; ChatGPT's free voice output is non-commercial and may not be repackaged as audio files (Service Terms §8) (openai.com service terms, updated Jun 12, 2026, checked Aug 11, 2026)
Cartesia Credits; vendor states 1 credit = 1 character for standard TTS (Pro Voice Clone TTS 1.5 credits/char) (cartesia.ai/pricing FAQ, checked Aug 11, 2026) Pro, $5/mo, 100K credits — includes commercial-use license and instant voice cloning; cheapest commercial-license entry point in this table (cartesia.ai/pricing, checked Aug 11, 2026) Pro: $5 ÷ 100,000 × 1,000 = $0.050/1k chars; Startup: $49 ÷ 1,250,000 × 1,000 = $0.039/1k; Scale: $299 ÷ 8,000,000 × 1,000 = $0.037/1k; overage rates $65/$45/$38 per 1M credits = $0.065/$0.045/$0.038 per 1k (cartesia.ai/pricing, checked Aug 11, 2026) 20K credits/mo (= 20,000 chars), $0. Expressly non-commercial — ToS restricts free use to “personal, non-commercial use only” (§4.1, §5.3(b)); card requirement not stated (cartesia.ai ToS, last revised Jun 14, 2024; pricing, both checked Aug 11, 2026)
MiniMax — API Characters (pay-as-you-go); subscriptions use “audio points” with no published points→characters conversion — subscriptions not normalizable (platform.minimax.io PAYG pricing, checked Aug 11, 2026) API Audio Subscription Starter, $5/mo, 100,000 audio points (value per character not computable); pay-as-you-go has no documented minimum (platform.minimax.io subscription pricing, checked Aug 11, 2026) speech-2.8-turbo: $60 / 1M ÷ 1,000 = $0.06/1k chars; speech-2.8-hd: $100 / 1M ÷ 1,000 = $0.10/1k chars. Rapid Voice Cloning $1.50 per voice; Voice Design $3 per voice (platform.minimax.io, checked Aug 11, 2026) No free API speech tier documented. Consumer app gives “daily time-limited free credits” (quota unpublished) but its ToS permits “personal, non-commercial use only” (paid service terms, updated Feb 12, 2026; consumer ToS, Jun 19, 2025, both checked Aug 11, 2026)
Resemble AI Seconds (billing-usage API reports seconds); subscription + prepaid credits — no characters conversion, and no public TTS rate at all (docs.resemble.ai billing usage, checked Aug 11, 2026) Team, $350/mo ($280/mo annual); Flex is $0/mo pay-per-use credits — but no voice-synthesis rate is published for either (resemble.ai/pricing, checked Aug 11, 2026) Not computable — TTS list pricing is not published. The public pricing page lists per-second rates only for the deepfake-detection products; no speech-synthesis or cloning rate appears (resemble.ai/pricing, checked Aug 11, 2026) Flex $0/mo, no card required, credits never expire; free voice-generation quota unpublished; ToS is silent on free-tier and paid-tier commercial rights (pricing; ToS, last update Jul 30, 2024, both checked Aug 11, 2026)

Feature matrix: API, cloning, languages, formats, licensing

The matrix below compares what each vendor documents, not what we have measured. Two columns deserve special care: language counts are the vendors' own claims — several vendors publish different counts on different pages, and the cells say so — and the commercial-use column covers paid tiers only; free-tier rights are in the pricing table above, because they are usually different and usually worse.

Feature comparison from official vendor documentation, checked August 11, 2026. Language counts are the vendors' own claims as of that date, not our measurements.
Provider API Voice cloning — consent / tier requirements Languages (vendor claim, Aug 11, 2026) Output formats Commercial use on paid tiers
ElevenLabs Yes — REST + WebSocket, Python/Node SDKs, per-request cost header (API docs, checked Aug 11, 2026) Instant cloning from Starter ($6/mo, per pricing page); Professional cloning Creator+ ($22/mo), own voice only, mic verification + 24 h wait; ToS requires the voice be yours or one you are authorized to share (cloning docs; ToS, checked Aug 11, 2026) Eleven v3: 70+; Flash v2.5: 32; Multilingual v2: 29; Flash v2: English only (models page, checked Aug 11, 2026) MP3, PCM, WAV, µ-law/A-law, Opus; 192 kbps needs Creator+, 44.1 kHz PCM/WAV needs Pro+; streaming via HTTP/WebSocket (API reference, checked Aug 11, 2026) Commercial license from Starter ($6/mo); user retains output rights (ToS §4(c), updated Mar 31, 2026) (elevenlabs.io/terms-of-use, checked Aug 11, 2026)
Google Cloud Text-to-Speech Yes — REST + gRPC, client libraries, bidirectional streaming, published per-project quotas (API reference, checked Aug 11, 2026) Chirp 3: Instant Custom Voice — sales allow-list only; fixed per-language consent script mandatory; reference audio up to 10 s; synthesis $60/1M chars (ICV docs, checked Aug 11, 2026) “380+ voices across 75+ languages and variants” (product page); Gemini-TTS: 84 languages listed (24 GA); Chirp 3 HD: 53 locales (product page and docs, checked Aug 11, 2026) LINEAR16 (WAV), MP3 (32 kbps), OGG_OPUS, MULAW, ALAW; model-dependent streaming sets; async long audio to 1M input bytes (AudioEncoding reference, checked Aug 11, 2026) B2B contract; customer retains IP in Customer Data (ToS §5.1) — no named “commercial use” grant needed or given; bans on competitive use and model training from output; no generative-output indemnity for TTS (cloud.google.com/terms; service terms, checked Aug 11, 2026)
Murf AI Yes — REST + HTTP streaming + WebSockets, official Python SDK, published rate limits (API docs, checked Aug 11, 2026) Enterprise-only, not self-serve; ~1–2 h studio-quality recordings, 1–4 week turnaround; explicit written consent required (ToS §4.2) (help.murf.ai; ToS, checked Aug 11, 2026) Murf's own pages disagree: 33–35+ languages, 150–300+ voices depending on page (murf.ai and help center, checked Aug 11, 2026) Studio export (paid only): MP3, WAV, FLAC, A-law, µ-law at 8/24/48 kHz (help.murf.ai); API: MP3, WAV, FLAC, ALAW, ULAW, PCM, OGG; streaming yes (API reference, both checked Aug 11, 2026) “You can use Murf created voices for commercial purposes” (ToS §5.1); no resale, no using output to train AI (ToS, Dec 19, 2024, checked Aug 11, 2026)
Speechify Yes — REST batch + streaming, Python/TS SDKs (docs.speechify.ai, checked Aug 11, 2026) Studio: paid plans only, 20-second sample; API: 10–30 s sample plus a mandatory consent record (speaker's full name + email); prohibited: minors, deceased persons, well-known political figures (cloning API docs; API terms, checked Aug 11, 2026) App/Studio: 60+ (marketing); API: simba-multilingual 30+, simba-3.0 six languages, simba-3.2 English only — different products, different counts (docs; pricing, checked Aug 11, 2026) API: mp3 and wav documented in examples (no full enum published); streaming endpoint to 20,000 chars/request; Studio download formats unpublished (API reference, checked Aug 11, 2026) Studio paid: explicit grant — commercial usage rights, “you own the audio output”; API: output is Customer Content with a mandatory AI-generated disclosure, but no express commercial grant; app Premium: non-commercial (ToS §4.4) (pricing-studio; ToS, checked Aug 11, 2026)
Amazon Polly Yes — REST + 9 SDKs + CLI, async batch to 100K billed chars, billed-character response header (API reference, checked Aug 11, 2026) No self-service cloning. Brand Voice is a negotiated enterprise engagement with the Polly team (actor recording; no published price/timeline). Responsible AI Policy forbids depicting a voice without consent (features page; policy, checked Aug 11, 2026) “100+ voices in 40+ languages and variants” (product page); docs voice list spans 42 language rows; generative engine: 43 voices (aws.amazon.com/polly; voice list, checked Aug 11, 2026) MP3, Ogg Vorbis, Ogg Opus, PCM, µ-law, A-law; Speech Marks JSON; streaming responses; bidirectional streaming for generative engine in 7 regions (API reference, checked Aug 11, 2026) “Your Polly output belongs to you”; “no restrictions on storing and reusing generated speech”; ownership not tier-scoped; only ban: training competing AI models (Service Terms §50.5) (FAQ; service terms, Jul 29, 2026, checked Aug 11, 2026)
Azure Speech in Foundry Tools Yes — REST + Speech SDK (C#/Python/JS/Java) + batch synthesis + containers; machine-readable price API (REST reference, checked Aug 11, 2026) Three routes, all Limited Access (“only customers managed by Microsoft”): personal voice (1-min sample), professional voice (30 min–3 h + biometric consent verification), MAI instant clone (5–60 s, preview). Recorded talent consent + disclosure duties mandatory (Limited Access page, checked Aug 11, 2026) “More than 500” prebuilt voices (HD voices comparison), 100+ languages/locales; personal voice speaks 91 languages; MAI-Voice-2: 15 languages (TTS overview and model docs, all checked Aug 11, 2026) MP3 (to 48 kHz/192 kbps), Opus/OGG/WebM, raw PCM, WAV/RIFF, A-law/µ-law, AMR-WB, G.722; streaming + batch; 10-min real-time cap per request (REST reference, checked Aug 11, 2026) Express grant, paid tier only: “For Customers of the paid tier TTS Service only, Customer may use the audio output of prebuilt neural voices … including for commercial purposes”; output = Customer Data; exclusive rights to custom voices (Product Terms, checked Aug 11, 2026)
OpenAI Yes — API only; no browser studio product; Python/JS SDKs; per-tier rate limits published (TTS guide, checked Aug 11, 2026) Yes, via API: consent recording of a fixed OpenAI phrase required first; samples ≤30 s; max 20 voices/org; a referenced “Text-to-Speech Supplemental Agreement” could not be located (TTS guide, checked Aug 11, 2026) 57 languages named (follows Whisper coverage); voices “optimized for English” (TTS guide, checked Aug 11, 2026) mp3 (default), opus, aac, flac, wav, pcm; chunked streaming; SSE stream format (not for tts-1/hd); 4,096-char request cap (API reference, checked Aug 11, 2026) Customer owns API output with express assignment (Services Agreement §4.1, eff. Jan 1, 2026); mandatory disclosure that the voice is AI-generated; ChatGPT Voice output non-commercial (business terms; service terms, checked Aug 11, 2026)
Cartesia Yes — REST bytes + SSE + WebSocket, JS/Python SDKs, published per-plan concurrency (docs.cartesia.ai, checked Aug 11, 2026) Instant clone (≤10 s clip) from Pro $5/mo; Pro Voice Clone: ≥30 min audio, Startup plan+ ($49/mo), 1M-credit training fee, output billed 1.5 credits/char; ToS bars cloning any person's voice without express permission; no consent-verification workflow documented (cloning docs; ToS, checked Aug 11, 2026) 42 languages (Sonic 3.5 model page, full code list published) (model docs, checked Aug 11, 2026) Containers raw/wav/mp3 (mp3+wav on Bytes endpoint only; SSE/WebSocket raw only); PCM f32le/s16le, µ-law, A-law; 8–48 kHz; MP3 32–192 kbps (format docs, checked Aug 11, 2026) “Commercial use license” from Pro ($5/mo); Cartesia claims no ownership of outputs (ToS §5.3); note: inputs/outputs train models by default, opt-out via form (§5.3(c)–(d)) (pricing; ToS, Jun 14, 2024, checked Aug 11, 2026)
MiniMax Yes — REST (sync to 10K chars, streaming), WebSocket, async long-form to 1M chars/request; billable-character counts in responses (API reference, checked Aug 11, 2026) Rapid Voice Cloning via API: $1.50/voice, 10 s–5 min sample; account verification gates cloning (error 2038); unused clones deleted after 7 days; customer bears data-subject consent duty; no consent-capture flow documented (cloning guide; platform ToS, eff. Mar 30, 2026, checked Aug 11, 2026) 40+ languages claimed; language_boost enum lists 40 named options; 332 system voices on the official ID list (API FAQ; voice list, checked Aug 11, 2026) mp3, pcm, flac, wav, G.711 µ-law, opus; 8–44.1 kHz (opus to 48); MP3 to 256 kbps; hex or 24-hour URL delivery; word/sentence timestamps (API reference, checked Aug 11, 2026) API/Open Platform ToS: customer retains ownership of generated content; no express commercial-use grant to quote and no restriction either; deep-synthesis disclosure marks required; consumer-app output (even paid) is not clearly commercially usable (platform ToS; consumer ToS, checked Aug 11, 2026)
Resemble AI Yes — sync REST + chunked HTTP streaming + WebSocket; timestamps; webhooks; SDK examples in 8 languages (API reference, checked Aug 11, 2026) Rapid Clone: 10 s–3 min audio, under 1 min training; Professional Clone: 10–25+ min, “explicit verifiable consent… required” (marketing FAQ); Voice Cloning API requires Business plan ($1,000/mo list); Rapid Clone consent verification ambiguous (cloning docs; product page, checked Aug 11, 2026) 100 languages claimed on the managed platform; 23 languages for open-source Chatterbox Multilingual (product page; GitHub, checked Aug 11, 2026) WAV (µ-law, PCM 16/24/32) + MP3; 8–48 kHz; HD flag; 3,000-char sync cap; PerTh watermark auto-embedded in generated audio (API reference; start-here page, checked Aug 11, 2026) ToS is silent-to-ambiguous: no output-ownership assignment and no commercial-use grant or restriction found; ban on using output to train/improve other products; open-source Chatterbox is MIT-licensed for self-hosting (ToS, last update Jul 30, 2024; GitHub, checked Aug 11, 2026)

ElevenLabs · expressive control at a premium

ElevenLabs is the most complete creator package in this comparison: the deepest documented expressive controls, self-serve voice cloning, and a commercial license that starts on the $6/month Starter plan (elevenlabs.io/pricing, checked Aug 11, 2026). The price reality: its API lists at $0.10 per 1,000 characters for Multilingual models (elevenlabs.io/pricing/api, checked Aug 11, 2026), and plan-effective rates run higher still. The bluntest con: it is the premium-priced option at every level — and its own two pricing pages disagree by roughly 1.65×–2×, depending on plan, on what a subscription includes.

What the documentation supports

The trade-offs

Cost, and cost transparency. Plan-effective rates work out to $0.165–$0.200 per 1,000 characters on the Multilingual basis (computed in the pricing table from elevenlabs.io/pricing, checked Aug 11, 2026) — roughly 41×–50× Google's and Amazon's $0.004 standard tiers, and about 11× Azure's $0.015 pay-as-you-go rate. Worse for budgeting: the main pricing page quotes plans in credits while the API pricing page quotes the same plans in characters at roughly 1.65×–2× the credit figures depending on plan, and the vendor does not reconcile the two (pricing vs API pricing, both checked Aug 11, 2026). Output quality is also tier-gated: 192 kbps MP3 requires Creator or above, and 44.1 kHz PCM/WAV requires Pro or above (API reference, checked Aug 11, 2026).

Free tier and commercial use

The free plan (10,000 credits/month) is strictly non-commercial and requires attribution (ToS §1(c), updated Mar 31, 2026; terms and billing docs, both checked Aug 11, 2026). Anything you intend to publish commercially needs at least Starter. We could not verify whether the free plan requires a card, or the exact attribution wording — the vendor's help article on it was inaccessible at check time (see what we could not verify).

Voice cloning and consent

Instant cloning is listed from the Starter plan; Professional Voice Cloning is restricted to your own voice, enforced by microphone verification and a 24-hour wait, and the ToS requires that any voice you clone be yours or one you are expressly authorized to share (cloning docs; terms, both checked Aug 11, 2026).

Who should look elsewhere

Cost-sensitive, high-volume API users: Google Cloud and Amazon Polly's standard tiers list at $0.004 per 1,000 characters (Google pricing; Polly pricing, both checked Aug 11, 2026) — 25× cheaper than ElevenLabs' Multilingual API rate. Anyone planning to publish work from a free tier, and anyone who needs to clone a voice other than their own through the professional route, should also pass. Instead consider Google Cloud Text-to-Speech for volume economics.

Google Cloud Text-to-Speech · the verified price floor

Google Cloud Text-to-Speech is the cheapest verified option in this comparison — $0.004 per 1,000 characters on the Standard and WaveNet tiers — and pairs it with the largest free allowance: 4 million characters per month on those tiers, always free (cloud.google.com/text-to-speech/pricing, checked Aug 11, 2026). The price reality: it is pay-as-you-go infrastructure with no plans and no studio. The bluntest con: even free usage requires a billing account, overage auto-bills, and voice cloning sits behind a sales allow-list.

What the documentation supports

The trade-offs

The newest models break the price comparison: Gemini-TTS is token-billed, and the only vendor conversion published (25 audio tokens per second) normalizes output to about $0.015/minute on 2.5 Flash and $0.030/minute on 2.5 Pro — with text-input tokens on top that cannot be converted at all (pricing, checked Aug 11, 2026). Studio voices cost $0.160 per 1,000 characters — 40× the Standard tier and above ElevenLabs' API list rate (same source). The terms ban competitive use and training models on the output, and TTS is not on Google's list of generative services covered by its output indemnity (service terms, checked Aug 11, 2026). Whether TTS falls under Google's separate "Generative AI Services" terms is genuinely ambiguous in the published documents — we say so rather than guess (see what we could not verify).

Free tier and commercial use

The free bands are generous but not casual: "You must enable billing to use Text-to-Speech", usage beyond the band auto-bills, and the $300/90-day trial requires a payment method (pricing; free-tier docs, both checked Aug 11, 2026). On commercial use, the terms draw no free-vs-paid distinction: all usage runs under the same B2B contract, in which the customer retains IP in Customer Data (ToS §5.1, cloud.google.com/terms, checked Aug 11, 2026).

Voice cloning and consent

Chirp 3: Instant Custom Voice is the cloning route, and it is gated: access is granted through a sales allow-list, a fixed per-language consent script must be recorded, reference audio is capped at 10 seconds, and synthesis with a custom voice costs $60 per 1M characters (ICV docs, checked Aug 11, 2026). There is no self-serve path.

Who should look elsewhere

Creators who want a browser studio, plan pricing, or self-serve cloning — none of which exist here. Buyers who need output IP indemnity for generated speech, and AI-adjacent companies constrained by the competitive-use ban, should read the service terms closely before committing. Instead consider Murf AI for an editor-first workflow with quotable commercial rights.

Murf AI · studio workflow, quotable commercial rights

Murf AI is the studio pick: a timeline-style editor sold as Studio plans from $29/month, next to a separately-priced API at $0.03 per 1,000 characters (murf.ai/pricing; API plans, both checked Aug 11, 2026) — and the cleanest commercial-use sentence in this comparison: "You can use Murf created voices for commercial purposes" (ToS §5.1, terms, checked Aug 11, 2026). The bluntest con: the free trial produces nothing you can keep — no downloads, self-evaluation only — and Studio time-billing cannot be compared per-character with rivals.

What the documentation supports

The trade-offs

Studio bills Voice Generation Time — hours of output — with no published character conversion, so its $0.133–$0.242/minute normalized range cannot be compared per-character against rivals; the vendor's own approximation is about 6 minutes of VGT per 1,000 words (help.murf.ai, checked Aug 11, 2026). Procurement teams should also know the public pricing page is a JavaScript-only shell: we verified prices from Murf's own production pricing endpoint, whose payload carries a "pricingTestGroup":"B" flag — meaning served prices may be A/B-test variants (murf.ai/pricing data, checked Aug 11, 2026). Murf's own pages also disagree on catalog size — 150 to 300+ voices and 33 to 35+ languages depending on the page (murf.ai and help center, checked Aug 11, 2026). One dated warning for API users: the legacy Gen2 streaming model deprecates on August 16, 2026 — five days after this page's verification date — and users must move to Falcon 2 (murf.ai/api/docs, checked Aug 11, 2026).

Free tier and commercial use

The Studio free trial gives 10 minutes of VGT with no card required — but no downloads of any kind, and the ToS limits trial use to "self-evaluation" (§10.7) (help.murf.ai; ToS, both checked Aug 11, 2026). It answers "should I subscribe?", not "can I use this output?". The API trial is more useful: 100,000 characters with one API key (help.murf.ai, checked Aug 11, 2026).

Voice cloning and consent

Cloning is enterprise-only and not self-serve: roughly 1–2 hours of studio-quality recordings, a 1–4 week turnaround, and explicit written consent from the voice owner required by ToS §4.2 (help.murf.ai; terms, both checked Aug 11, 2026).

Who should look elsewhere

Anyone who wants self-serve voice cloning — Murf has none at any self-serve tier. Anyone who evaluates tools through free output will find the trial useless by design. And buyers who need vendor-stable public pricing pages for procurement should note the A/B-tested, JavaScript-only pricing surface. Instead consider ElevenLabs, where instant cloning is listed from the $6/month tier.

Speechify · the cheapest API plans, with a licensing trap

Speechify's developer platform is the cheapest entry-level character-billed subscription here: the API Starter plan is $10/month for 1 million characters — $0.010 per 1,000 — falling to $0.006 per 1,000 in Scale-tier overage, a rate that ties Azure's largest commitment tier (speechify.ai/pricing, checked Aug 11, 2026). The price reality: "Speechify" is three products with three different sets of terms. The bluntest con: the $29/month Premium app most consumers find first is expressly non-commercial (ToS §4.4, speechify.com/terms, checked Aug 11, 2026), and the API's commercial rights rest on inference, not on a quotable grant.

What the documentation supports

The trade-offs

The product split does real damage to clarity. The consumer app's Premium plan is metered in words (a 150,000-word/month contractual floor, with 1M words/month guaranteed through 2026 — usage limits, checked Aug 11, 2026) and its terms list "commercial use" among the grounds for restricting an account (speechify.com/usage-limits, checked Aug 11, 2026; consistent with the non-commercial rule in ToS §4.4, terms, checked Aug 11, 2026). The API's supplemental terms designate output as "Customer Content" and impose a duty to disclose AI generation (§1.4) — but contain no express commercial-use grant to quote (API terms, Jun 10, 2024, checked Aug 11, 2026). Studio displays only annual prices, and its download formats are unpublished (pricing-studio, checked Aug 11, 2026).

Free tier and commercial use

Three free tiers, three answers. The API free tier (50K chars/month) sits under terms that are silent on free-vs-paid commercial rights — we state that silence rather than fill it (API terms, checked Aug 11, 2026). Studio's free plan states outright: "No commercial usage rights", and no voice cloning (pricing-studio, checked Aug 11, 2026). The app's free tier offers 10 "robotic sounding voices" and is non-commercial like Premium (pricing, checked Aug 11, 2026).

Voice cloning and consent

Studio cloning requires a paid plan and a 20-second sample. API cloning takes a 10–30 second sample plus a mandatory consent record containing the speaker's full name and email, and prohibits cloning minors, deceased persons, and well-known political figures (cloning API docs; API terms, both checked Aug 11, 2026).

Who should look elsewhere

Anyone about to buy Premium expecting to publish commercial work — that is the single easiest licensing mistake in this comparison. API customers whose legal review requires an express commercial grant in writing, and multilingual API projects needing current-generation quality beyond simba-3.0's six languages, should also look on. Instead consider Murf AI, whose paid-tier commercial grant is one quotable sentence.

Amazon Polly · lowest-cost infrastructure with the cleanest ownership

Amazon Polly matches the cheapest verified rate in this comparison — $0.004 per 1,000 characters on the Standard engine, pay-as-you-go with no plans (aws.amazon.com/polly/pricing, checked Aug 11, 2026) — and pairs it with the cleanest ownership language: "Your Polly output belongs to you", with "no restrictions on storing and reusing generated speech" (Polly FAQ, checked Aug 11, 2026). The bluntest con: no voice cloning, no creator UI — and AWS's own free-tier pages contradict each other.

What the documentation supports

The trade-offs

The newest quality tier comes with caveats AWS itself documents: the Generative engine is available in only 9 regions, produces no Speech Marks, and carries a vendor-acknowledged hallucination risk with an "emergency stop mechanism" that "reduces but does not eliminate the risk" (developer guide, checked Aug 11, 2026). By default AWS may store and use processed text to improve its AI services, with opt-out only via an AWS Organizations policy — worth knowing for confidential scripts (Service Terms §50.3, service terms, checked Aug 11, 2026). And there is no self-serve cloning of any kind (features page, checked Aug 11, 2026).

Free tier and commercial use

The free tier is two-track and ambiguous. Per-engine monthly character bands exist (Standard 5M; Neural 1M; Long-Form 500K; Generative 100K — the latter three "for the first 12 months", while the Standard band's duration is stated inconsistently across AWS's own pages), but accounts opened on or after July 15, 2025 fall under a $200-credit regime whose precedence over the character bands AWS does not state (pricing; free-tier FAQ, both checked Aug 11, 2026). A card is required at signup regardless. On commercial use, no tier-scoped restriction appears anywhere — free-band output runs under the same service terms as paid (service terms, checked Aug 11, 2026).

Voice cloning and consent

There is no self-service cloning. Brand Voice — a custom voice built with the Polly team from actor recordings — is a negotiated enterprise engagement with no published price or timeline, and AWS's Responsible AI Policy forbids depicting a person's voice without consent (features; policy, both checked Aug 11, 2026).

Who should look elsewhere

Anyone who wants voice cloning, a consumer editor, or output beyond the AWS console-and-SDK workflow. Teams outside the 9 Generative-engine regions who want its newest tier, and testers unwilling to hand AWS a card just to experiment, should also pass. Instead consider ElevenLabs for cloning and an editor, or Google Cloud Text-to-Speech for the same price floor with bigger free bands.

Azure Speech in Foundry Tools · the biggest catalog, behind enterprise gates

Azure Speech in Foundry Tools — Microsoft's renamed Azure AI Speech (Learn TTS overview, checked Aug 11, 2026) — claims the largest catalog among the cloud platforms in this comparison: "more than 500" prebuilt voices (HD voices comparison, checked Aug 11, 2026) across 100+ languages and locales (Learn TTS overview, same date). The price reality: $0.015 per 1,000 characters pay-as-you-go, falling to $0.006 on the largest commitment tier (Retail Prices API, checked Aug 11, 2026). The bluntest con: free-tier output sits outside the commercial-use grant, and every voice-cloning route requires being "managed by Microsoft".

What the documentation supports

The trade-offs

Transparency, oddly, is the weak spot for a company this size: the canonical pricing page rendered every dollar figure as "$-" placeholders at check time, so our figures come from Microsoft's official Retail Prices API (eastus region — prices vary by region), corroborated by the Learn documentation's $15/1M rate (prices.azure.com; Learn quotas doc, both checked Aug 11, 2026). The newest models — MAI-Voice-2 and MAI-Voice-2-Flash — are public preview, carry no SLA, and Microsoft says they are not recommended for production (model docs via the TTS overview, checked Aug 11, 2026). Two billing quirks: every CJK character counts as two billed characters (same overview), and signup requires a non-prepaid card for identity verification even for free use (account page, checked Aug 11, 2026).

Free tier and commercial use

The F0 tier gives 0.5M neural characters per month at no cost — but read the grant: commercial use of prebuilt neural voice output is licensed "For Customers of the paid tier TTS Service only", which leaves F0 output outside the quoted permission (Microsoft Product Terms, checked Aug 11, 2026). Treat F0 as an evaluation tier, not a production one.

Voice cloning and consent

All three cloning routes — personal voice (1-minute sample), professional voice (30 minutes to 3 hours of recordings), and the MAI instant clone preview (5–60 seconds) — are Limited Access features: "only customers managed by Microsoft" are eligible, recorded talent consent is required and may be verified biometrically, and AI-disclosure duties to end users are mandatory (Limited Access page, checked Aug 11, 2026).

Who should look elsewhere

Hobbyists planning to monetize on the free tier — the grant does not cover them. Anyone who wants self-serve cloning without an enterprise relationship, and teams that need the newest voice models under a production SLA, should also pass. Instead consider Cartesia, where a commercial license and instant cloning are self-serve at the lowest verified entry price.

OpenAI · steerable speech with no free way in

OpenAI's Audio API is the developer's shortcut to instruction-steerable speech: gpt-4o-mini-tts takes plain-language direction over accent, emotion, tone, and even whispering (TTS guide, checked Aug 11, 2026). The price reality: legacy tts-1 lists at $0.015 and tts-1-hd at $0.030 per 1,000 characters, but the current flagship model is token-billed with no published conversion (pricing, checked Aug 11, 2026). The bluntest con: there is no free way to try it, and no way to budget its flagship model per character from official documentation.

What the documentation supports

The trade-offs

The catalog is small and English-leaning: 13 built-in voices (9 on the legacy models), described by the vendor as "optimized for English", with 57 languages claimed via Whisper lineage (TTS guide, checked Aug 11, 2026). Requests cap at 4,096 characters (2,000 tokens for gpt-4o-mini-tts), forcing chunking for long-form (API reference, checked Aug 11, 2026). gpt-4o-mini-tts bills $0.60/1M text-input tokens and $12/1M audio-output tokens with no token-to-character or token-to-minute conversion published, so its real-world cost cannot be normalized (pricing, checked Aug 11, 2026). There is no browser studio product at all. Apps must disclose to end users that the voice is AI-generated (TTS guide, checked Aug 11, 2026).

Free tier and commercial use

No free TTS tier of any kind exists; API use requires prepaid credits purchased in advance, which expire after one year and are non-refundable — a card is effectively required before any use (prepaid billing help, checked Aug 11, 2026). ChatGPT's free voice feature is not a workaround: its output is expressly non-commercial and may not be repackaged as standalone audio files (Service Terms §8, updated Jun 12, 2026, service terms, checked Aug 11, 2026). Paid API output, by contrast, is owned by the customer outright under the §4.1 assignment quoted above.

Voice cloning and consent

Custom voices arrived via the API: a consent recording of a fixed OpenAI-specified phrase is required first, samples are capped at 30 seconds, and organizations may hold at most 20 voices (TTS guide, checked Aug 11, 2026). One honesty note: the documentation references a "Text-to-Speech Supplemental Agreement" that we could not locate anywhere — cloning-rights specifics may live in a document we could not read (see what we could not verify).

Who should look elsewhere

Anyone who needs a free tier, a browser UI, or a large voice library. Long-form producers unwilling to chunk at 4,096 characters, and buyers who must budget the flagship model's cost precisely in advance, should also pass. Instead consider Google Cloud Text-to-Speech, which pairs steerable Gemini-TTS models with free bands and a published character price floor.

Cartesia · budget real-time voice with self-serve cloning

Cartesia has the cheapest commercial entry point we verified: the $5/month Pro plan includes 100,000 credits, a commercial-use license, and instant voice cloning from a clip of 10 seconds or less (cartesia.ai/pricing, checked Aug 11, 2026). The price reality: a published 1-credit-per-character conversion makes its rates fully comparable — $0.037–$0.050 per 1,000 characters by plan (same source). The bluntest con: by default your inputs and outputs may train Cartesia's models unless you file an opt-out, and the governing terms predate the plans they govern.

What the documentation supports

The trade-offs

The data-governance default is the big one: the ToS states that inputs, outputs, and interactions "may be used by Cartesia to train… its models", with opt-out by request form (§5.3(c)–(d), terms, checked Aug 11, 2026). The same ToS is dated June 14, 2024 — before the current plan names existed — so the commercial-tier mapping rests on the pricing page rather than the contract (terms; pricing, both checked Aug 11, 2026). Format support is also asymmetric: MP3 and WAV containers are available only on the Bytes endpoint, while the SSE and WebSocket streaming endpoints return raw audio only (format docs, checked Aug 11, 2026).

Free tier and commercial use

The free tier gives 20,000 credits (20,000 characters) per month and is expressly non-commercial: the ToS restricts free use to "personal, non-commercial use only" (§4.1, §5.3(b), terms, checked Aug 11, 2026). Whether a card is required to sign up is not stated (pricing, checked Aug 11, 2026).

Voice cloning and consent

Instant cloning from a ≤10-second clip is included from Pro; the higher-fidelity Pro Voice Clone requires at least 30 minutes of audio, a Startup plan or above ($49/month), a one-time 1M-credit training fee, and bills output at 1.5 credits per character (cloning docs; pricing, both checked Aug 11, 2026). The ToS bars making any person's voice available without express permission (§4.2(13)) — but no consent-verification workflow is documented in the cloning docs, so enforcement rests on the contract, not the product (terms, checked Aug 11, 2026).

Who should look elsewhere

Teams with strict data-governance requirements who cannot rely on an opt-out form, MP3-dependent workflows built on streaming endpoints, and anyone who needs consent-verified cloning as a product feature rather than a contract clause. Instead consider Azure Speech in Foundry Tools, whose enterprise terms make data handling and voice exclusivity explicit.

MiniMax · cheap, capable, and legally split

MiniMax's speech API is aggressively priced: speech-2.8-turbo at $0.06 and speech-2.8-hd at $0.10 per 1,000 characters pay-as-you-go, with Rapid Voice Cloning at $1.50 per voice — the lowest documented cloning fee in this comparison (platform.minimax.io pricing, checked Aug 11, 2026). The bluntest con: the licensing surface is split and partly silent — the consumer app's output is not clearly commercially usable even on paid plans, and the API terms grant ownership without ever saying "commercial use permitted".

What the documentation supports

The trade-offs

Subscriptions are denominated in "audio points" with no published points-to-characters conversion, so their value cannot be cost-compared — only pay-as-you-go can (subscription pricing, checked Aug 11, 2026). No free API speech tier is documented, so there is no documented way to evaluate before paying (pricing, checked Aug 11, 2026). The API contract is with Nanonoble Pte. Ltd. under the platform ToS effective March 30, 2026, which requires deep-synthesis disclosure — services built on it must add identifiers and "a prominent mark" telling the public deep synthesis was used (platform ToS, checked Aug 11, 2026).

Free tier and commercial use

The consumer MiniMax Audio app offers "daily time-limited free credits" with an unpublished quota, but its ToS permits "personal, non-commercial use only" (consumer ToS, Jun 19, 2025; paid service terms, updated Feb 12, 2026, both checked Aug 11, 2026) — and the paid consumer plans' commercial position is ambiguous. On the API side, the platform ToS states the customer retains ownership of generated content, with no non-commercial restriction — but also no affirmative commercial-use grant to quote (platform ToS, checked Aug 11, 2026). Commercial users belong on the API, and cautious ones should ask MiniMax for the grant in writing.

Voice cloning and consent

Rapid Voice Cloning costs $1.50 per voice plus synthesis-priced previews, from a 10-second to 5-minute sample (pricing, checked Aug 11, 2026). Cloning is gated on account verification, unused cloned voices are deleted after 7 days, and the data-subject consent duty sits with the customer — no consent-capture flow is documented in the product (cloning guide; platform ToS, both checked Aug 11, 2026).

Who should look elsewhere

Consumer-app users who need commercial rights, buyers who need a documented free tier to evaluate against, and procurement teams that require an express commercial grant or convertible subscription units. Instead consider OpenAI, whose terms assign output ownership in one unambiguous sentence.

Resemble AI · security-first voice, price on request

Resemble AI is the security-and-consent specialist of this list: every generated file carries the PerTh perceptual watermark by vendor statement, consent tooling is built into professional cloning, and the MIT-licensed Chatterbox family is the standout self-hosting route in this comparison (start-here page; GitHub, both checked Aug 11, 2026). The bluntest con: you cannot learn what its hosted voice generation costs from any public page, and its terms never say who owns the audio or whether you may sell it.

What the documentation supports

The trade-offs

Pricing opacity is unmatched in this comparison — in the wrong direction. The hosted API's only current TTS model is resemble-ultra; every previous hosted model, including all hosted Chatterbox variants, is end-of-life and "can no longer generate audio" (model-versions doc, checked Aug 11, 2026). Yet no voice-generation rate is published anywhere: the public pricing page lists per-second rates only for the deepfake-detection products, and the app's pricing URL returned a 404 at check time (resemble.ai/pricing, checked Aug 11, 2026). Plans are listed at Team $350/month and Business $1,000/month with no synthesis rate attached (same page). The company's public packaging now centers deepfake detection and AI security; voice generation remains fully documented but visibly de-emphasized (pricing; product page, both checked Aug 11, 2026).

Free tier and commercial use

Flex is $0/month with no card and never-expiring credits, but the free voice-generation quota is unpublished (pricing, checked Aug 11, 2026). More importantly, the ToS contains no output-ownership assignment and no commercial-use grant or restriction at any tier — a silence a buyer must resolve with sales before committing; the one clear prohibition is using output to train or improve other products (ToS, last update Jul 30, 2024, checked Aug 11, 2026). Self-hosted Chatterbox is different: MIT-licensed, so its permissions are in the license itself (GitHub, checked Aug 11, 2026).

Voice cloning and consent

Rapid Clone builds a voice from 10 seconds to 3 minutes of audio in under a minute of training; Professional Clone takes 10–25+ minutes of audio and, per the vendor's FAQ, requires "explicit verifiable consent" with consent workflows "built into the platform" — though how Rapid Clone verifies consent is ambiguous in the docs (cloning docs; product page, both checked Aug 11, 2026). Note the gate: the Voice Cloning API requires the Business plan at $1,000/month list (same sources).

Who should look elsewhere

Anyone who needs transparent self-serve pricing, cloning-API access below a $1,000/month plan, or explicit output-rights language their legal team can quote. Instead consider Cartesia, which publishes every rate and sells a commercial license from $5/month.

Who should choose what

These recommendations are derived from the documented facts above — pricing, licensing, and feature coverage as of August 11, 2026 — not from listening tests. Each links back to the full assessment.

Video and content creators

ElevenLabs is the default: expressive tags, self-serve cloning, and a commercial license from $6/month (pricing, checked Aug 11, 2026) — if the per-character premium fits your volume. Teams that live in a timeline editor should weigh Murf AI, whose paid plans pair the workflow with the plainest commercial-use sentence in this comparison. Speechify Studio paid plans state output ownership explicitly, but display annual prices only.

Developers and API builders

For volume, Google Cloud Text-to-Speech ($0.004/1k characters and 4M free characters/month on Standard/WaveNet — pricing, checked Aug 11, 2026) and Amazon Polly (same floor, cleanest ownership terms) are the anchors. For subscription budgeting, Speechify's API is the cheapest character-billed plan class we verified; Murf AI's API at $0.03/1k with a $2 minimum (API plans, checked Aug 11, 2026) is the low-friction middle. Teams already on OpenAI infrastructure get instruction steering and clean ownership from OpenAI — priced per character only on its legacy models.

Enterprise and compliance-led buyers

Azure Speech in Foundry Tools is built for this buyer: commitment tiers to $0.006/1k characters, on-prem containers, customer-exclusive custom voices, and explicit data terms (Retail Prices API; Product Terms, both checked Aug 11, 2026). Google Cloud offers comparable contractual footing with a stated no-training commitment. Where audio provenance itself is the requirement, Resemble AI's default watermarking is the documented differentiator.

Real-time agents and conversational apps

Latency numbers here are all vendor claims — none are our measurements. Cartesia claims sub-90 ms to first audio and publishes per-plan concurrency (pricing, checked Aug 11, 2026); ElevenLabs claims ~75 ms model latency on Flash v2.5 (models page, checked Aug 11, 2026); Murf AI claims sub-130 ms time to first audio on Falcon (murf.ai/api, checked Aug 11, 2026). All three ship WebSocket streaming; MiniMax's turbo tier is the budget entrant. When our benchmark runs, this is one of the first claims it will test.

Budget and free usage

The largest documented free allowances are Google Cloud's 4M characters/month (Standard/WaveNet) and Amazon Polly's 5M Standard characters/month band (Google pricing; Polly pricing, both checked Aug 11, 2026) — both under B2B terms with no tier-scoped commercial restriction found, but both requiring billing setup. Speechify's API free tier (50K chars/month) is the rare hard-capped one. The caution: most creator-tool free tiers are expressly non-commercial — see the FAQ. The cheapest paid commercial licenses we verified are Cartesia at $5/month and ElevenLabs at $6/month (Cartesia pricing; ElevenLabs pricing, both checked Aug 11, 2026).

Self-hosted and on-device

One documented route exists in this comparison: Resemble AI's MIT-licensed Chatterbox family — Turbo (350M), Nano (110M, on-device), and Multilingual V3 (23+ languages), usable without an API key (GitHub, checked Aug 11, 2026). Azure Speech offers on-prem containers, but as a licensed cloud service rather than open source.

What we could not verify

Honesty about gaps beats false completeness. The items below are questions our source records explicitly mark as unverifiable from official documentation as of August 11, 2026 — the page states them rather than guessing. Where a vendor's own pages disagree with each other, both figures are attributed to their pages above, never averaged.

Frequently asked questions

What is the cheapest AI voice generator?

By verified list price: Google Cloud Text-to-Speech and Amazon Polly, both at $0.004 per 1,000 characters on their standard tiers, pay-as-you-go (Google pricing; Polly pricing, both checked Aug 11, 2026). Among subscription APIs, Speechify's Starter plan works out to $0.010 per 1,000 characters (speechify.ai/pricing, checked Aug 11, 2026). But "cheapest" depends on the tier that matches your quality needs — the normalized table shows the full spread.

Can I use a free AI voice generator commercially?

Usually not. Explicitly non-commercial or excluded from the commercial grant as of August 11, 2026: ElevenLabs' free plan (ToS §1(c)), Cartesia's free tier (ToS §4.1), Speechify Studio's free plan ("No commercial usage rights"), Azure's F0 tier (paid-tier-only grant), Murf's trial (self-evaluation only, §10.7), and the MiniMax consumer app (personal, non-commercial only). The exceptions: AWS and Google Cloud terms contain no tier-scoped restriction — free-band usage runs under the same B2B contracts as paid (AWS service terms; Google Cloud terms, both checked Aug 11, 2026) — and OpenAI has no free TTS at all.

Which free tier is actually the largest?

By documented volume: Amazon Polly's Standard band (5M characters/month, though AWS states its duration inconsistently) and Google Cloud's always-free 4M characters/month on Standard/WaveNet voices (Polly pricing; Google pricing, both checked Aug 11, 2026). Both require billing setup (AWS requires a valid payment method at signup; Google requires an enabled billing account). Among no-cloud options, Speechify's API gives 50,000 characters/month with a hard cap that pauses rather than bills (speechify.ai/pricing, checked Aug 11, 2026).

Why can't some prices be compared directly?

Because vendors bill in incompatible units and do not always publish conversions. OpenAI's gpt-4o-mini-tts and Google's Gemini-TTS are token-billed with no published token-to-character conversion (OpenAI pricing; Google pricing, both checked Aug 11, 2026); Murf Studio sells hours of "Voice Generation Time" (help.murf.ai, checked Aug 11, 2026); MiniMax subscriptions use unconvertible "audio points" (platform docs, checked Aug 11, 2026). Where no vendor conversion exists, our table says "not normalizable" instead of inventing one.

What happened to PlayHT?

We can only report what we observed: on August 11, 2026, play.ht resolved to no DNS A record, and neither did play.ai; both domains' authoritative DNS now sits on Meta nameservers. We could not verify a cause from any official source, so this page does not assert one — PlayHT is simply not listable while its site is unreachable and its pricing and terms cannot be checked.

Do I have to tell listeners a voice is AI-generated?

Increasingly, yes — by contract if not by law. OpenAI requires apps to disclose that the voice is AI-generated (TTS guide); Speechify's API terms require "commercially reasonable disclosures" of AI generation (§1.4, API terms); MiniMax requires deep-synthesis identifiers and "a prominent mark" (platform ToS); Azure imposes disclosure duties for custom voices (Limited Access page); and ElevenLabs' free plan requires attribution (billing docs; non-commercial per ToS §1(c)). All checked Aug 11, 2026. This is a legal-obligations pattern, not legal advice.

Which providers offer self-serve voice cloning?

Cheaply and without a sales call, as of August 11, 2026: Cartesia (from $5/month, ≤10 s clip — pricing), ElevenLabs (instant cloning listed from $6/month — pricing), MiniMax ($1.50 per voice via API, gated on account verification — pricing), and Speechify (Studio paid plans and the API, with a mandatory consent record — docs). OpenAI's API cloning requires a fixed consent phrase (TTS guide). Enterprise-gated entirely: Murf, Azure, Google, Amazon (no self-serve cloning at all), and Resemble's cloning API (Business plan, $1,000/month list). Every route carries consent obligations — see each provider's cloning section above.

Do these tools train AI models on my scripts or audio?

It varies more than any other term we checked. Cartesia may train on inputs and outputs by default, with opt-out via a request form (§5.3(c)–(d), terms); AWS may store and use processed text to improve its services, with opt-out via AWS Organizations policy (§50.3, service terms); Google states it will not train on customer data without permission (§18, service terms). All checked Aug 11, 2026. If your scripts are confidential, this clause belongs at the top of your checklist.

Did you listen to these voices?

No — and this page says so wherever it matters. Our standardized audio benchmark (same short prompts, same conditions, every provider) is designed but not yet running; see the methodology. Everything above comes from official vendor documentation checked on August 11, 2026. When the benchmark runs, this page will be refreshed with the results and the change log below will record it.

Change log

Published August 11, 2026 · all prices and terms read from official sources on this date.