ElevenLabs Review 2026
Published August 11, 2026 · Every fact verified against ElevenLabs' official documentation the same day.
Everything this site has verified about ElevenLabs — this review, the alternatives, every head-to-head, its pricing position and the contradictions in its own published material — is indexed on the ElevenLabs hub.
Quick verdict
ElevenLabs is the most complete self-serve AI voice package we have documented — expressive current-generation models, voice cloning from $6/month, an explicit commercial license on every paid plan, and a serious developer API. It is our #1 pick among ten providers in the market-wide ranking. The cost of that completeness is twofold: premium pricing, and a pricing structure that ElevenLabs' own two pages describe differently. Verdict dated Aug 11, 2026.
| Best for | Creators and product teams who want expressive voices, self-serve cloning and clean commercial licensing in one subscription |
|---|---|
| Not best for | Cost-driven high-volume TTS (Google and Amazon list 25× cheaper); publishing from a free plan; cloning voices other than your own |
| Starting price | $6/month Starter (30,000 credits); API from $0.05–$0.10 per 1,000 characters (pricing, API pricing) |
| Commercial use | Allowed on every paid plan; prohibited on the free plan (ToS §1(c), updated Mar 31, 2026) |
| Test status | Documentation-verified. Our listening benchmark has not run, so there is no independent voice-quality score yet. |
We verified ElevenLabs' features, pricing, API documentation and commercial-use terms against the company's official pages, checked August 11, 2026. Our standardized listening benchmark has not yet been completed, so we do not assign an independent voice-quality score; vendor claims about how the voices sound are labeled as theirs throughout.
Strongest advantages:
- Commercial license from $6/month, with output rights retained by you (ToS §1(c), §4(c)).
- Self-serve voice cloning at consumer prices — instant from Starter, professional from Creator.
- The deepest documented expressive controls in our coverage: v3 audio tags, native IPA pronunciation, five tunable voice settings.
- A transparent developer API: official Python and Node.js SDKs, WebSocket streaming, a per-request cost header.
Important disadvantages:
- Premium pricing: plan-effective cost computes to $0.165–$0.200 per 1,000 Multilingual-model characters.
- Two official pricing pages describe plan inclusions on different bases, disagreeing by ~1.65×–2× without reconciliation.
- Free plan: 10,000 credits/month, strictly non-commercial, title attribution required.
- Top audio quality is paywalled: 192 kbps MP3 needs Creator+; 44.1 kHz PCM/WAV needs Pro+.
On this page
Who should use it — and who should look elsewhere
Choose ElevenLabs if:
- You produce character-driven narration — YouTube, storytelling, ads. The v3 model accepts inline direction like
[whispers]and[sarcastic], which no other self-serve product in our coverage documents at this depth. - You want to clone your own voice without a sales call. Instant cloning starts at $6/month on 1–2 minutes of audio; the professional route adds microphone-verified ownership from $22/month.
- Your legal review needs one quotable sentence. Paid subscribers “may use the Services for commercial purposes” (ToS §1(c), updated Mar 31, 2026), and you retain rights in what you generate (§4(c)).
- You are building TTS into a product. REST and WebSocket APIs, official SDKs, and a
character-costheader on every response make billing transparent per generation.
Look elsewhere if:
- Price per character drives the decision. Google Cloud and Amazon Polly list at $0.004 per 1,000 characters on standard tiers — 25× below ElevenLabs' $0.10 Multilingual API rate (Google pricing; Polly pricing, checked Aug 11, 2026). See the normalized table.
- You hoped to publish commercially from a free plan. ElevenLabs' free tier is expressly non-commercial, and published free-plan output requires title attribution.
- You need to clone a client's or talent's voice professionally. Professional Voice Cloning is own-voice-only — even with consent.
- Your budgeting can't absorb ambiguous accounting. ElevenLabs' credits-based and characters-based pricing pages disagree on what a subscription includes, and the vendor doesn't reconcile them.
Current price, free plan, commercial rights
The short version: commercial use starts at $6/month, the free plan is a demo, and the API bills separately at $0.05–$0.10 per 1,000 characters. Full breakdown in pricing and normalized cost.
| Entry price | Free at $0 (10,000 credits/month); first paid plan $6/month (30,000 credits) |
|---|---|
| API list price | $0.05/1k characters (Flash/Turbo); $0.10/1k characters (Multilingual v2 / v3) |
| What plans really cost | $0.165–$0.200 per 1,000 Multilingual-model characters, plan-effective (computation below) |
| Free plan commercial use | No — “you may only use the Services for non-commercial purposes” (ToS §1(c)); attribution required when publishing |
| Paid plan commercial use | Yes — “Commercial License” from Starter up; output rights stay with you (§4(c)) |
| Annual billing | Pay for 10 months, get 12: equivalent to $5 Starter / $18.33 Creator / $82.50 Pro / $249.17 Scale / $825 Business |
| Unused credits | Roll over up to two months on active paid plans (balance so the balance can reach at most 3× the monthly quota); no rollover on Free |
What this means for you: if you need commercial output, budget for at least Starter — there is no free-tier grace period for publishing. If your volume is low and API-only works, pay-as-you-go at the list price undercuts every subscription's effective rate (details below). And because the two pricing pages disagree on plan inclusions, budget against the conservative credits basis.
What this review covers
Four current text-to-speech models, documented as of Aug 11, 2026 (models docs):
| Model (API id) | Role, per ElevenLabs | Languages (vendor claim) | Max characters per generation |
|---|---|---|---|
Eleven v3 (eleven_v3) |
The expressive flagship — “human-like and expressive”, “most emotionally rich” | 70+ | 5,000 |
Eleven Multilingual v2 (eleven_multilingual_v2) |
The consistent long-form model — “most lifelike… rich emotional expression”; the API default | 29 | 10,000 |
Eleven Flash v2.5 (eleven_flash_v2_5) |
The speed-and-cost tier — “ultra-fast… real-time use (~75ms†)” (vendor claim) | 32 | 40,000 |
Eleven Flash v2 (eleven_flash_v2) |
English-only speed tier; the only model with SSML phoneme tags | English only | 30,000 |
Deprecated models (eleven_turbo_v2_5, eleven_turbo_v2, eleven_multilingual_v1) are excluded. The practical split: v3 for expression, Multilingual v2 for stability, Flash v2.5 for real-time economics. Voice selection — stock library, designed voices, clones — is covered under cloning and workflow; we do not evaluate individual stock voices, which would require hands-on access we don't claim.
Our benchmark standard
When our benchmark runs, ElevenLabs will be measured against the same three short prompts we use for every provider (full design on the methodology page):
| Test | What it probes | Prompt |
|---|---|---|
| A. Naturalness | Pacing, phrasing, pauses, realism | “A good voice should feel effortless: clear enough to follow, warm enough to trust, and natural enough that you stop thinking about the technology behind it.” |
| B. Precision | Numbers, dates, currency, codes | “Your booking is Friday, August 21st at 7:45 p.m. The total is $129.50, and your confirmation code is A7X4.” |
| C. Expression (optional) | Energy, emphasis, emotional control | “We finally made it. After months of work, the doors are open, the lights are on, and tonight everything begins.” |
Test C is the interesting one here: v3's audio tags are built for exactly this kind of direction. That's a note about the design, not a result.
Benchmark audio status
Samples are pending. No audio has been generated, so no player appears here — we don't publish empty placeholders. In place of listening results, this review verifies what documentation can establish (all checked Aug 11, 2026): the model line and per-model limits (models docs), the full output-format list and quality tier gates (API reference), the documented control surface (prompting docs), and the pricing and licensing documents linked throughout. When benchmark audio exists, it will appear here with model, voice, duration and actual cost.
Benchmark cost status
No measured benchmark cost exists yet. The only cost figures on this page are documented prices, each with source and date (summary, full breakdown). For planning only: at the documented $0.10 per 1,000 Multilingual-model characters (API pricing, checked Aug 11, 2026), our ~400-character suite would cost on the order of four cents — an estimate from list prices, not a measurement.
Assessment by dimension
No numeric scores yet — those wait on the benchmark and a finalized, published methodology (how that works). For now, a documentation-verified read of each dimension:
| Dimension | Where ElevenLabs stands on paper |
|---|---|
| Naturalness | Open question — vendor claims “human-like” and “most lifelike” for v3 and Multilingual v2. Awaiting our listening benchmark. |
| Pronunciation | Strongest documented toolkit in our coverage: native IPA in v3 (vendor-stated 80–90% consistency), SSML phonemes in Flash v2, PLS dictionaries — but fragmented across models. |
| Expression / control | Deepest documented controls we've found: v3 audio tags (laughter, whispers, sarcasm, accents, sound effects), five tunable settings. |
| Consistency | Multilingual v2 is positioned as the “lifelike, consistent quality” model with the largest quality-tier cap (10,000 chars). Awaiting listening benchmark. |
| Voice cloning | Two self-serve routes with consent checks built in: instant (Starter+) and professional (Creator+, own-voice-only, microphone-verified). |
| Languages | Vendor claims: 70+ (v3), 32 (Flash v2.5), 29 (Multilingual v2). Claims only — no per-language evaluation exists on this site yet. |
| Speed | Vendor claims ~75 ms for Flash v2.5. We publish measurements only under documented conditions; none exist yet. |
| Editor / UX | Browser studio with projects, Productions and dubbing on one credit pool. No hands-on observations — no product access granted or used. |
| API & integrations | Strong: REST + WebSocket, official SDKs, cost headers, machine-readable formats. |
| Price / value | The weak spot: premium list prices, plan-effective $0.165–$0.200/1k, and two unreconciled pricing pages. |
Voice quality and naturalness
Documentation can't tell you whether a voice sounds human, so we won't pretend it can. What it can tell you is the quality ceiling — and ElevenLabs gates that ceiling by plan.
The API reference enumerates output from telephony grade up to 44.1 kHz, with two explicit gates: “MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above” and “PCM and WAV formats with 44.1kHz sample rate requires you to be subscribed to Pro tier or above” (API reference, checked Aug 11, 2026). Free and Starter top out at 128 kbps per the pricing compare table. If your deliverable is broadcast-quality audio, the plan you need is decided before you press generate.
ElevenLabs' own quality claims, attributed: v3 is “human-like and expressive” and the vendor's “most emotionally rich” model; Multilingual v2 is “our most lifelike model”; Flash v2.5 is the real-time tier while “Multilingual v2 delivers the highest quality audio with more nuanced expression” (models docs; capabilities docs, checked Aug 11, 2026). Whether those claims survive our benchmark prompts is exactly what the listening test will settle.
Pronunciation, numbers, difficult text
Misread prices, dates and confirmation codes are the most common way production TTS embarrasses you, so control here matters more than marketing. ElevenLabs documents three mechanisms (prompting docs, checked Aug 11, 2026):
- Native IPA in v3 — transcriptions in forward slashes across “70+ languages”, with a rare honest caveat: “V3's IPA support achieves 80-90% pronunciation consistency… it is not 100% consistent.”
- SSML phoneme tags in Flash v2 only — CMU Arpabet or IPA; the docs are explicit that Multilingual v2 “doesn't support phoneme tags”.
- PLS pronunciation dictionaries — TXT or .PLS uploads, usable in the Studio, Dubbing Studio and API; matching is case-sensitive, first-match-wins.
Pauses work differently by generation: v2 models accept SSML breaks up to 3 seconds, while “Eleven v3 does not support SSML break tags” and relies on punctuation and audio tags (same source).
What this means for you: the deepest pronunciation tooling currently sits on the cheap Flash v2 model and on v3's native IPA — not on Multilingual v2, the default. If your content is heavy in brand names and codes, pick the model with the control you need, and plan for our precision prompt (“$129.50”, “A7X4”) to test it properly once the benchmark runs.
Expression, pacing, control
This is ElevenLabs' documented signature strength. Five voice settings are tunable per request (API reference, checked Aug 11, 2026):
| Setting | Default | Role |
|---|---|---|
stability | 0.5 | Steadiness vs variability of delivery |
similarity_boost | 0.75 | Closeness to the source voice |
style | 0 | Style exaggeration |
use_speaker_boost | true | Speaker boost |
speed | 1.0 | 0.7–1.2 range (prompting docs) |
The headline feature is v3-only: audio tags. Inline directives — [laughs], [whispers], [sighs], [sarcastic], [curious], [crying], [sings], [strong X accent], even sound effects like [applause] — direct the performance from inside the script. The docs caution that “some experimental tags may be less consistent across different voices” (same source/date).
What this means for you: no other self-serve product in our coverage documents this granularity of performance direction. It is documented capability, not demonstrated output — the listening benchmark will say whether a [sarcastic] tag produces sarcasm you'd actually ship.
Voice cloning
Two routes, both self-serve, both built around consent (cloning docs, checked Aug 11, 2026):
| Instant Voice Cloning | Professional Voice Cloning | |
|---|---|---|
| Audio needed | 1–2 minutes of good audio; ~30 s can work; beyond 2–3 min “will yield little improvement” | 30–180 minutes of good audio |
| Cheapest plan | Starter, $6/month (pricing page; billing docs: “Cloning is only available on the Starter tier and above”) | Creator, $22/month (“available on our Creator plan or above”) |
| Whose voice | Yours, or one “you are authorized to share” (ToS §4(b)); cloning “without consent or legal right” is prohibited | Own voice only — “Even with their consent, you cannot clone someone else's voice.” Owners can share a verified clone privately via link. |
| Verification | None described | Mandatory microphone check; failed attempts wait 24 hours |
| Turnaround | Instant | Fine-tuning “generally… 3-6 hours”, up to 24 |
| Slots | Not slot-limited in the docs read | Creator/Pro: 1 · Scale: 3 · Business: 10 · Enterprise: custom; extras via “Studio Quality” review |
| Portability | Clones stay on ElevenLabs — no export. Downgrade below Creator and a PVC stays in your library but unusable. | |
Source audio guidance: MP3 at 192 kbps or above, one speaker, clean recording. The compliance picture is buyer-relevant: the ToS puts authorization on you, the use policy bans non-consensual replication and deception about AI origin, and generated audio “can be instantly traced back to the user responsible”.
What this means for you: creators cloning their own voice get the cheapest documented self-serve route in our ranking. Agencies wanting to clone client talent hit a hard wall on the professional route — by design. The flagship's feature matrix sets this against every competitor's cloning regime.
The professional tier carries the strictest restriction in this market, and it is worth reading before you plan around it: “You can only create a Professional Voice Clone of your own voice. Even with their consent, you cannot clone someone else’s voice.” Sample floors are 1 minute for instant and 30 minutes for professional, with fine-tuning stated at 3–6 hours depending on queue. Our guide to how AI voice cloning works sets those against the other nine.
ElevenLabs also publishes the clearest account in this market of what a cloned voice actually is: “Voice cloning captures a representation of a voice, not a recording of it”, and instant cloning “uses your audio sample as a conditioning signal at inference time… without any model weight updates”. That is the real mechanical difference between the two tiers. Our guide to how AI voice generators work quotes it in full.
Languages and accents
Vendor claims, dated — not evaluated (models docs, checked Aug 11, 2026):
| Model | Claimed coverage |
|---|---|
| Eleven v3 | 70+ languages, with native IPA pronunciation control across them |
| Eleven Flash v2.5 | 32 — “all languages from v2 models plus: Hungarian, Norwegian & Vietnamese” |
| Eleven Multilingual v2 | 29 languages |
| Eleven Flash v2 | English only |
Accents can be directed in v3 via tags like [strong X accent] (prompting docs, checked Aug 11, 2026). Per-language quality testing doesn't exist anywhere on this site yet, so treat the counts as coverage claims, not quality claims — check the models page for your specific language, and weigh it as a vendor statement until multilingual benchmarking exists.
One page, two incompatible figures: the heading on the text-to-speech page says “over 70 languages” while the FAQ lower on that same page says “32+ languages across our model lineup”. The two enumerated lists disagree as well — we counted 74 entries in the docs and 76 names in the marketing list. Neither is endorsed here; both are published in our multilingual comparison.
Editor, workflow, export
What documentation establishes: one browser app bundles text to speech, speech-to-text, sound effects, voice design, music, image, Productions and Studio projects — 3 projects on Free, 20 from Starter, with Dubbing Studio listed from Starter (pricing, checked Aug 11, 2026). Export quality follows the tier gates above; for long text, the capabilities docs advise segmenting with streaming playback.
What this review can't tell you: how the studio actually feels — editing, project flow, failure modes. We haven't used it, and we won't describe hands-on experience we don't have. If we ever collect that evidence under the stored-artifact rule in our editorial policy, it will be labeled first-hand.
API, SDK, integrations
A developer surface that respects your billing team (API docs; API reference, checked Aug 11, 2026):
- Transport: HTTP or WebSocket “from any language”; API-key auth.
- SDKs: official Python (
pip install elevenlabs) and Node.js (npm install @elevenlabs/elevenlabs-js). - Core endpoint:
POST /v1/text-to-speech/{voice_id}; default modeleleven_multilingual_v2; model list viaGET /v1/models. - Cost transparency: every response carries
character-cost,request-idandx-trace-idheaders — per-generation cost logging without estimation. - Output control: machine-readable formats across MP3, PCM, WAV, Opus and telephony codecs, tier gates as noted above.
- Billing separation: “API usage is billed in US dollars, not credits” (API pricing) — with the plan-inclusion caveat in the disagreement table.
The same account also reaches Scribe speech-to-text ($0.22/hour list), Music ($0.15/min), Sound Effects, Voice Isolator, Voice Changer and Dubbing — relevant if your product needs more than TTS. Rate limits per plan are not published on any discoverable docs page (unverified list).
Two cautions for integrators. The large SDK list belongs to the Agents Platform, not the REST text-to-speech API, which has two — and the libraries most likely to appear in a JavaScript or game-engine stack sit in a table marked as not officially supported. The streaming endpoint also accepts 21 output formats against the non-streaming endpoint’s 28, with no wav_* format at all. Our text-to-speech API comparison covers both.
Generation speed and latency
No independent measurements — we publish latency only under documented, repeatable conditions, and the benchmark isn't running. ElevenLabs' claim, attributed: Flash v2.5 is “optimized for real-time use (~75ms†)” (models docs; capabilities docs, checked Aug 11, 2026), and the Business plan advertises “Low-latency TTS as low as 5c/minute” (pricing). The dagger in the vendor's own copy signals conditions apply; the pages read don't define them. Treat every number here as a vendor claim.
ElevenLabs publishes the most careful latency documentation in this market and it undercuts its own headline. Its concepts page states that model inference latency “is an internal measurement, excluding network round-trips and application overhead”, while time-to-first-audio “is almost always the number that matters for user experience, and it is always larger”. Its own regional table gives 100–150 ms TTFB across three regions and 150–200 ms across two more. Our real-time comparison sets all three of its figures side by side.
Pricing and normalized cost
ElevenLabs' pricing is easy to read and harder to compare — because the company describes what you get on two pages that don't agree. Here is the official price, what you actually receive, what it works out to, and who gets the best value. All figures verified Aug 11, 2026.
Subscriptions: what you pay, what you get
| Plan | Price/month | Credits/month | What you actually get |
|---|---|---|---|
| Free | $0 | 10,000 | TTS, speech-to-text, sound effects, voice design, music, Productions, image, 3 projects. No commercial license, no cloning, no rollover. |
| Starter | $6 | 30,000 | Commercial license, instant voice cloning, 20 projects, commercial music use, Dubbing Studio, image & video |
| Creator | $22 (first month $11) | 121,000 | Professional Voice Cloning, extra credits; the “Popular” tier; unlocks 192 kbps MP3 |
| Pro | $99 | 600,000 | 44.1 kHz PCM via API, 192 kbps audio quality |
| Scale | $299 | 1,800,000 | 3 seats, team collaboration, 3 professional clones |
| Business | $990 | 6,000,000 | “Low-latency TTS as low as 5c/minute”, 10 professional clones, 10 seats |
| Enterprise | Custom | Custom | DPA/SLAs, HIPAA BAAs, custom SSO, elevated concurrency, managed dubbing, volume discounts |
Pay annually and you pay for 10 months out of 12 — equivalent to $5 Starter, $18.33 Creator, $82.50 Pro, $249.17 Scale, $825 Business (same page/date). Unused credits roll over up to two months while a paid plan stays active.
Pay-as-you-go API: the clean numbers
| Item | List price |
|---|---|
| TTS — Flash / Turbo models | $0.05 per 1,000 characters (also rendered “~$0.05/minute”) |
| TTS — Multilingual v2 / v3 models | $0.10 per 1,000 characters (also rendered “~$0.10/minute”) |
| Scribe v2 / Realtime speech-to-text | $0.22 / $0.39 per hour |
| Music / Sound Effects | $0.15 / $0.12 per minute (effects billed per generation) |
| Dubbing v1 / v2 | $0.33/min with watermark ($0.50 without) / $2.20/min |
| Pay-as-you-go terms | “No commitment, cancel anytime”, all products and models; no separate PAYG rates shown |
Effective normalized cost
Subscriptions look cheap until you divide price by included volume. Using the pricing page's own credit basis (1 credit = 1 Multilingual character, per its FAQ):
| Plan | Computation | Effective $/1k chars | vs its own API list ($0.10) |
|---|---|---|---|
| Starter $6 | 6 ÷ 30,000 × 1,000 | $0.200 | 2.0× |
| Creator $22 | 22 ÷ 121,000 × 1,000 | $0.182 | 1.8× |
| Pro $99 | 99 ÷ 600,000 × 1,000 | $0.165 | 1.7× |
| Scale $299 | 299 ÷ 1,800,000 × 1,000 | $0.166 | 1.7× |
| Business $990 | 990 ÷ 6,000,000 × 1,000 | $0.165 | 1.7× |
Context from the ranking's normalized table (checked Aug 11, 2026): Google Cloud's standard tiers and Amazon Polly list at $0.004/1k, Azure at $0.015, Speechify's entry API plan at $0.010. ElevenLabs' effective band is the premium end of that spread. What you're paying for is everything above the price: expression, cloning, the studio, the API ergonomics.
The important limitation: two pages, two stories
| Plan | Main pricing page (credits basis) | API pricing page (characters basis) | Ratio |
|---|---|---|---|
| Starter $6 | 30,000 credits (≈30,000 Multilingual chars) | 60,000 Multilingual / 120,000 Flash chars | ≈2.0× |
| Creator $22 | 121,000 credits | 220,000 / 440,000 chars | ≈1.8× |
| Pro $99 | 600,000 credits | 990,000 / 1,980,000 chars | ≈1.65× |
| Scale $299 | 1,800,000 credits | 2,990,000 / 5,980,000 chars | ≈1.66× |
| Business $990 | 6,000,000 credits | 9,900,000 / 19,800,000 chars | ≈1.65× |
The API page's figures equal plan price ÷ list price (Creator: $22 ÷ $0.10/1k = 220,000). The main page's credits feed a shared pool where TTS costs 1 credit per Multilingual character. Billing docs note that credit consumption depends on “whether you're generating via the website or API” — consistent with two metering bases, but ElevenLabs never reconciles the pages. We won't solve that contradiction by guessing. Budget against the conservative credits basis.
One more number worth knowing: the pricing page's compare table shows UI overage at roughly $0.36/min (Free), $0.20 (Starter), $0.18 (Creator), $0.17 (Pro and up), against included UI minutes of ~10 to ~6,000 across the plans (same page/date; the vendor renders these with “~”).
Who gets the best value
- Light, API-only commercial usage: pay-as-you-go at $0.10/1k beats every plan's effective rate — cheapest documented commercial route if the studio doesn't matter to you.
- Creators who want cloning + license: Starter at $6 is ElevenLabs’ own lowest commercial entry point, though Cartesia’s $5 Pro plan is cheaper still across this comparison; you pay a 2× markup over list for the bundled products.
- High-volume production: Pro and above flatten the effective rate to ~$0.165/1k — still premium, but the per-character penalty stops growing.
- Annual subscribers: two months free; Starter effectively $5/month.
At one credit per character on V2 Multilingual, the Starter plan works out to $200 per million characters if you use the whole allowance — fifty times the lowest published rate in this market, and half that again on Flash. The full normalized table is in our pricing comparison.
Dubbing is metered separately from text-to-speech and priced per source audio minute rather than per character — $0.33 watermarked, $0.50 without, and $2.20 on the alpha v2 model. ElevenLabs is the only one of the four vendors selling dubbing that publishes a rate you can read on a page; our AI dubbing comparison sets it against the other three.
Commercial usage, licensing, privacy
The free plan cannot be used commercially. Every paid plan can. That is the entire tier structure, quoted from the current Terms of Service (non-EEA, updated March 31, 2026; terms, checked Aug 11, 2026):
ToS §1(c): “if you access or use our Services free of charge (such a user, a ‘Free User’), you may only use the Services for non-commercial purposes; if you access or use our Services through a paid subscription plan (such a user, a ‘Paid User’), you may use the Services for commercial purposes, but in either case, your access and use of the Services and any Output must still comply with the Prohibited Use Policy.” Non-commercial use also requires a personal e-mail address (same section). The pricing page lists “Commercial License” from Starter up. Rights are stated per tier exactly as written — we don't extrapolate paid-tier rights onto the free tier.
Free-plan attribution, verbatim
If you publish free-plan output even non-commercially: “you must attribute it to ElevenLabs by including ‘elevenlabs.io’ or ‘11.ai’ in the title” (Eleven Music output: reference “Eleven Music”) — official help-center legal article, checked Aug 11, 2026. Same article: paid plans include the commercial license “provided you're not using Beta Services”, and Beta-Services output “cannot be used for any commercial purpose or in any production environment”. Billing docs add the plain version: “If you are on the free plan, you can use the content non-commercially with attribution.”
Who owns what you generate
You do. ToS §4(c): “as between you and ElevenLabs, you retain all rights in and to your Input” and “you retain all rights in and to your Output” — with the clarification that Output doesn't include the underlying models. ElevenLabs keeps the models; you keep the audio.
Cloning rights and consent
§4(b) allows User Voice Models from “your voice or the voice you are authorized to share with us”, deletable through your account. The Prohibited Use Policy (updated Aug 17, 2026) bans replicating a voice “without consent or legal right” and any use “intended to deceive others about whether the voice was generated by artificial intelligence” (use policy). Professional cloning enforces own-voice-only at the product level (see voice cloning).
Privacy and open questions
Generated audio “can be instantly traced back to the user responsible for the generation” — relevant provenance language for compliance teams. The Voice Processing Notice lives in ElevenLabs' Privacy Policy, which was outside this review's scope. Two questions stay open: whether the free plan requires a payment method, and ElevenLabs' data-training terms for inputs and outputs — both recorded as unverified below rather than guessed.
ElevenLabs is one of three providers that grant commercial use in an explicit, quotable clause rather than leaving it to inference. Our guide to can you use AI voices commercially compares that grant with the nine others.
Pros
- One subscription covers the whole job — TTS, speech-to-text, dubbing, sound effects, music and voice design on one credit pool (pricing, checked Aug 11, 2026).
- Expressive control with documentation behind it — v3 audio tags, five tunable settings, native IPA (vendor-stated 80–90% consistency).
- Commercial license from $6/month; you keep rights in your output (§1(c), §4(c)).
- Self-serve cloning in two tiers of seriousness — instant from Starter, microphone-verified professional from Creator.
- Billing-transparent API — per-request
character-costheader, official SDKs, WebSocket streaming. - Current model line, honest deprecations — v3 flagship, Multilingual v2 for stability, Flash v2.5 for real-time economics.
- Unused credits roll over up to two months on active paid plans.
Cons
- Premium price at every level — $0.10/1k API list; $0.165–$0.200/1k plan-effective; up to 50× Google's and Amazon's standard tiers.
- Two pricing pages that disagree by ~1.65×–2× on what a subscription includes, unreconciled by the vendor.
- The free plan is a demo — 10,000 credits, non-commercial, attribution required, no cloning, no rollover.
- Quality is paywalled — 192 kbps MP3 from Creator, 44.1 kHz PCM/WAV from Pro.
- Pronunciation tooling fragmented across models — phoneme SSML only in Flash v2, native IPA only in v3, neither in the default Multilingual v2.
- Professional cloning is own-voice-only — a hard wall for agencies cloning client talent.
- Rate limits unpublished on any discoverable docs page; concurrency terms appear only as an Enterprise bullet.
- Per-generation character caps — 5,000 on flagship v3 — force chunking for long-form work.
Best use cases
Fit by use case, judged from the evidence above — editorial judgment, not test results:
| Use case | Fit | Why |
|---|---|---|
| Narration with character — YouTube, storytelling, ads | Strong | v3 audio tags and style controls; commercial license from $6 |
| Cloning your own voice | Strong | Instant from Starter on 1–2 minutes of audio; professional route adds verification |
| Real-time / conversational agents | Strong on paper | Flash v2.5: vendor-claimed ~75 ms, $0.05/1k list, 40,000-char cap |
| Long-form narration — audiobooks, courses | Good | Multilingual v2's 10,000-char cap and consistency positioning — at premium per-minute cost |
| Multilingual catalogs | Good, claims only | 70+ languages claimed on v3; no per-language evaluation exists yet |
| High-volume, cost-led API TTS | Weak | Plan-effective $0.165–$0.200/1k vs $0.004–$0.016 at Google, Amazon, Azure standard tiers |
| Publishing on $0 | Weak | Free tier is non-commercial with attribution; no rollover |
Two use cases where ElevenLabs’ documented capability is specifically relevant. It is one of only three providers here that can produce multi-speaker dialogue in a single request, through a dedicated endpoint taking up to 10 unique voices. And for book-length work it positions Multilingual v2 as “Most stable on long-form generations” — though its async endpoint returns MP3 only, which matters if your distributor wants lossless masters.
Alternatives
Full alternatives page: ElevenLabs alternatives — the same question worked through in full, organised by reason for leaving: cheaper per character, self-serve cloning, developer tooling, multilingual coverage, a usable free tier, enterprise deployment and long-form work, with a section on when ElevenLabs is still the right choice.
Organized by the reason you'd leave — every target live in our ranking (assessments at the linked sections):
- Cheapest verified per-character cost: Google Cloud Text-to-Speech ($0.004/1k Standard/WaveNet, 4M free chars/month) and Amazon Polly ($0.004/1k standard engines).
- Cheapest commercial entry with cloning: Cartesia — $5/month Pro bundles license and instant cloning from as little as 10 seconds of audio.
- Studio workflow, quotable one-sentence commercial grant: Murf AI.
- Cheapest entry-level API subscription: Speechify — $10/month for 1M API characters.
- Instruction-steerable speech inside the OpenAI stack: OpenAI.
- Enterprise scale with compliance surface: Azure Speech in Foundry Tools.
All checked Aug 11, 2026, and all priced and licensed in the normalized table.
Direct comparison links
Two head-to-head comparisons naming ElevenLabs are now published, both documentation-based: ElevenLabs vs Google Cloud Text-to-Speech — the cheapest self-serve cloning against the lowest verified per-character rate — and ElevenLabs vs Murf AI — cloning access against a studio workflow. Both were built from documented facts checked September 1, 2026 and carry no listening-test claims; when our audio benchmark runs they gain a real head-to-head test. The flagship’s pricing, feature matrix and who should choose what sections remain the ten-provider view.
Final verdict
Choose ElevenLabs when the voice is the product: creators, studios and product teams whose output lives or dies on expression, and who want cloning, licensing and an API from one vendor without a sales call. For that buyer, no product we have documented packages more of the job into a $6–$22/month subscription — and the licensing is clean enough to quote to a lawyer.
Don't choose it when the voice is a commodity. If you're converting millions of characters where any competent voice will do, Google and Amazon list at a fortieth of the effective price, and ElevenLabs' premium buys you nothing you'll use. The same goes for anyone who needs to clone voices other than their own at the professional tier, and anyone whose finance team can't tolerate a vendor whose two pricing pages tell different stories about what a plan includes.
What would change this verdict is sound. Everything above is documented capability and quoted terms; the question this review can't answer yet is whether v3's audio tags produce performances worth the premium. When our benchmark runs, this page gains the samples and measured costs to settle it — and this verdict will be revisited on evidence, not assumption.
Frequently asked questions
Is ElevenLabs free?
There is a free plan — 10,000 credits/month at $0 — but it's a trial in everything but name: non-commercial only, title attribution required when publishing, no cloning, no credit rollover (ToS §1(c), updated Mar 31, 2026; pricing, checked Aug 11, 2026). Commercial work starts at $6/month.
Can I use ElevenLabs commercially?
Yes, on any paid plan — “if you access or use our Services through a paid subscription plan… you may use the Services for commercial purposes” (ToS §1(c), updated Mar 31, 2026), and the pricing page lists a Commercial License from Starter up. Two caveats: free-plan output stays non-commercial, and Beta-Services output is excluded from commercial use per the help-center legal article. Full quotes in licensing.
Which plan is best for commercial use?
For most creators, Starter ($6/month) — it's the cheapest plan with the commercial license and instant cloning. Heavy voice cloning needs Creator ($22) for the professional route. If you only need API volume and no studio, pay-as-you-go at $0.10/1k Multilingual characters undercuts every plan's effective rate ($0.165–$0.200/1k) — see who gets the best value.
How much does ElevenLabs actually cost per 1,000 characters?
API list: $0.10 (Multilingual v2/v3) or $0.05 (Flash/Turbo). Subscriptions work out to $0.165–$0.200 per 1,000 Multilingual-model characters depending on plan, computed from plan price ÷ included credits — and note the vendor's two pricing pages disagree on included volume by ~1.65×–2× (table). All verified Aug 11, 2026.
Can I clone someone else's voice?
Professional cloning: no — own voice only, “even with their consent”, enforced by microphone verification. Instant cloning: the voice must be yours or one “you are authorized to share” (ToS §4(b)), and cloning without consent or legal right is prohibited. Voice owners can clone on their own account and share privately via link. Details in voice cloning.
Who owns the audio I generate?
You do — “you retain all rights in and to your Output” (ToS §4(c), updated Mar 31, 2026) — with the standing condition that free-plan output stays non-commercial. Voice clones themselves can't be exported; they live on ElevenLabs.
Which ElevenLabs model should I use?
Per the vendor's documentation (checked Aug 11, 2026): v3 for maximum expressiveness (audio tags, 70+ claimed languages, 5,000-char cap); Multilingual v2 — the API default — for consistent long-form (10,000-char cap); Flash v2.5 for real-time and cost (vendor-claimed ~75 ms, half the Multilingual price, 40,000-char cap). Need phoneme-level pronunciation control? That currently means Flash v2 or v3's native IPA — not Multilingual v2.
Is ElevenLabs better than the alternatives?
For expressive, self-serve, license-clean creation — on documentation, it's our #1 of ten (ranking). For cost-led volume, Google Cloud and Amazon Polly are 25× cheaper at list; for the cheapest commercial cloning entry, Cartesia at $5. “Better” depends on which column drives your decision — the alternatives section is organized by reason.
Did you actually test ElevenLabs' voices?
Not yet. This review verifies what official documentation establishes — features, pricing, terms, limits — and says so everywhere; quality claims are attributed to ElevenLabs, not adopted. Our standardized listening benchmark is designed but not yet running (methodology); when it runs, samples and measured costs appear here with a dated change-log entry.
Sources and what we could not verify
Every changing fact on this page was read from official ElevenLabs sources on August 11, 2026. The complete research record — verbatim quotes, computations, unresolved disagreements — is kept in this site's research dataset for this review.
| Source | Used for |
|---|---|
| elevenlabs.io/pricing | Plans, credits, feature lists, annual billing, rollover, overage, credits-per-character FAQ |
| elevenlabs.io/pricing/api | API list prices, included characters per plan, pay-as-you-go and enterprise terms |
| elevenlabs.io/docs/models | Model line, model IDs, language counts, character limits, deprecations |
| elevenlabs.io/terms-of-use | ToS (non-EEA, updated Mar 31, 2026): §1(c), §4(b), §4(c) |
| elevenlabs.io/use-policy | Prohibited Use Policy (updated Aug 17, 2026): consent, deception, free-user commercial use |
| TTS capabilities docs | Formats, streaming, quality positioning |
| Voice cloning docs | Instant/professional cloning requirements, tiers, verification, slots, own-voice rule |
| TTS API reference | Endpoint, default model, formats, voice settings, quality tier gates |
| API introduction | Auth, SDKs, transport, usage headers |
| Billing docs | Credit mechanics, free-tier assignment, cloning tier gate, attribution, rollover |
| Prompting/controls docs | Speed, audio tags, SSML rules, IPA, PLS dictionaries |
| Help-center legal article | Free-plan status, attribution format, paid-plan license, Beta Services caveat |
What we could not verify
Genuinely unresolved, recorded rather than guessed (Aug 11, 2026):
- Whether the free plan requires a payment method — billing docs list accepted methods but are silent on the requirement.
- The exact Flash/Turbo credit rate on subscriptions — published only as “between 0.5 and 1 credit per character”.
- A numeric API rate limit — still undocumented. The per-plan concurrency ladder was located on September 2, 2026 and is published in our API comparison; a requests-per-minute threshold is not.
- The EEA Terms of Service variant — all ToS quotes here are the non-EEA document.
- Reconciliation of credits-based vs characters-based plan inclusions — both official pages recorded and attributed; ElevenLabs does not reconcile them.
- Data-training terms for inputs/outputs — not extracted this session; no claim made either way.
Change log
2026-09-01 — Correction. This page cited the Prohibited Use Policy as “updated Sep 3, 2025”. Re-checked September 1, 2026, that document now displays “Last Updated 17 August 2026” — six days after this review’s original verification date. The citation has been corrected in both the body and the sources table. The policy is incorporated into the commercial-use rule by Terms of Service clause 1(c), so its version date is material. We have not re-read the full policy text against our August summary of it; that re-read is scheduled and any resulting change will be recorded here. All other facts on this page remain as verified August 11, 2026.
August 11, 2026 — Editorial rewrite: same verified facts, restructured for decision-first reading; consolidated benchmark-status statements; no factual changes.
Published August 11, 2026 · all prices, terms and feature statements read from official sources on this date. Benchmark audio, measured costs and scores will be added with a dated entry here when the audio benchmark runs (methodology).