Best AI Voice Cloning Software in 2026
Published September 2, 2026 · Every price, access gate and consent requirement below was read from the vendor’s own documentation on September 1, 2026 and carries that check date. This page ranks documented capability and controls, not how any cloned voice sounds — no listening test has been run.
Status correction, September 2, 2026. (correction record) Resemble AI appears in the ranking below and no longer sells voice cloning. The vendor names it directly: it does not sell “text-to-speech, voice cloning, speech-to-speech conversion, or voice design as commercial products”. Its open-source models remain published, and the vendor now calls them “research artifacts rather than supported products” carrying “no SLA, paid tier, or enterprise support”. Resemble AI’s own machine-readable description, dated August 13, 2026 and read on September 2, 2026, states verbatim: “Resemble AI does not sell voice generation products and is not accepting new voice customers. Voice continues as open research only.” The full review carries the correction notice and the surrounding quotations.
Quick verdict
Cloning access in this market runs from $5 a month to a $1,000-a-month plan to “contact sales” — and price tracks almost nothing else. Cartesia at $5/month and ElevenLabs at $6/month are the cheapest self-serve routes. Resemble AI documents its cloning API as requiring a $1,000/month plan. Amazon Polly, Azure and OpenAI publish no self-serve route at all.
The axis that should decide this, and usually does not, is consent. Some providers verify it by recording before a voice can exist; others take a warranty from you and check nothing. Those are very different products wearing the same label, and the difference is not reflected in the price.
The finding a buyer most needs: the cheapest documented cloning route in this comparison — MiniMax’s pay-as-you-go rapid cloning, with no subscription required — is also the one where we found no voice-owner consent requirement published anywhere in its developer documentation. That is a searched-and-not-found result across both verification passes, not an oversight.
Ranking basis: documented capability, access and consent controls only, checked September 1, 2026 — no accounts created, no voices cloned, no listening test, no numeric scores (how we verify).
The ranking
The order weighs three documented properties: whether cloning is reachable without a sales conversation, what it costs, and how seriously consent is enforced. Enforcement is weighted heavily — a cheap cloning tool with no consent gate is not a better product, it is a riskier one.
| # | Provider | Cheapest documented route | Consent enforcement |
|---|---|---|---|
| 1 | ElevenLabs | $6/month Starter (instant); $22/month Creator (professional) | Strong. Professional cloning is your own voice only, with identity verification |
| 2 | Speechify | $10/month API Starter | Strongest technical enforcement documented. Machine-verified by recording; no unattended path exists |
| 3 | Cartesia | $5/month Pro (instant); $49/month Startup (professional) | Contractual only. A warranty you give; the cloning API has no consent field |
| 4 | MiniMax | Pay-as-you-go, no subscription — the cheapest documented route here | None found. Identity-gated on the developer; no voice-owner requirement published |
| 5 | Resemble AI | Cannot be reconciled: a changelog says “first clone free, then $2 each”; the docs require the $1,000/month Business plan | Discretionary. Resemble “may require consent”, and consent “needs to be verbal” |
| 6 | Google Cloud Text-to-Speech | No plan; $60 per 1M characters, but access is allow-listed via sales | Mandatory and scripted. A recorded consent statement submitted with the audio |
| 7 | Azure Speech in Foundry Tools | Cannot be priced. Limited Access, Microsoft-managed customers only | Most heavily enforced here. Mandatory recorded statement, backed by biometric matching |
| 8 | OpenAI | No price published. Custom voices limited to eligible customers via sales | Structurally enforced. Two separate recordings required |
| 9 | Murf AI | Not self-serve at any tier — an enterprise engagement | Strong on paper. Explicit written consent from every speaker, with a defined term |
| 10 | Amazon Polly | No price published. Brand Voice is a sales-led custom engagement | Weak as documented — and nothing found is Polly-specific |
On this page
How we ranked, and what we do not claim
Three documented properties, weighted in this order: consent enforcement, whether the route is self-serve, and price.
Consent leads deliberately, and the reason is not squeamishness. A cloning product that verifies the speaker consented is a materially different thing from one that asks you to promise it did — different in legal exposure, different in what happens when a voice is misused, and different in what your own customers can do with it. Ranking on price alone would put the provider with no documented consent requirement near the top, which would be a bad recommendation dressed as an objective one.
What this ranking is not: it is not a quality ranking. Our standardized listening benchmark is designed but not yet running, so nothing here says which cloned voices sound most like their source. That is precisely the question a naturalness benchmark answers, and we have not run one. No numeric scores appear anywhere on this page.
Consent: the axis that should decide this
Grouped by what the provider actually does, rather than what it says it values.
Verified before a voice can exist
Speechify has the strongest technical enforcement we documented anywhere. Verbatim: “Consent is verified, not asserted. You never send a checkbox or a signed form: the speaker reads a phrase Speechify issues, and the recording of them doing so is checked and retained as the consent record for the voice.” The obvious workaround is closed explicitly: “Batch or unattended creation. There is no consent path without a speaker present.”
Azure Speech enforces a mandatory recorded statement for both of its cloning tiers, backed by biometric matching against the training audio. That is the most heavily enforced regime in this comparison — and it sits behind an eligibility gate that excludes anyone without a Microsoft account team.
OpenAI requires two separate recordings to create a custom voice, structurally enforcing the consent step rather than accepting an attestation.
Google Cloud requires a recorded, scripted consent statement submitted alongside the training audio.
Restricted by whose voice you may clone
ElevenLabs takes a different route to the same protection: professional cloning is restricted to your own voice, with identity verification. Verbatim: “You can only create a Professional Voice Clone of your own voice. Even with their consent, you cannot clone someone else’s voice.” That is narrower than a consent process — it forecloses third-party cloning at the professional tier entirely.
Murf AI requires explicit written consent from every speaker in the training audio, with a defined term, as part of an enterprise engagement rather than a self-serve flow.
A warranty you give, checked by nobody
Cartesia places the obligation on you contractually — its Acceptable Use Policy requires that you submit “your own voice and audio recordings or those of others with explicit consent” — while its Clone Voice API contains no occurrence of the word “consent” and no attestation field. The request takes a clip, a name, a language and metadata.
Resemble AI markets consent tooling, but its Terms are discretionary and set a verbal standard: “Resemble may require consent form the individual or third party whose voice is being cloned. Consent needs to be verbal, unless otherwise stated by Resemble.” The typographical error is the vendor’s own. Its product page describes something stricter; where a contract and a product page disagree, the contract governs.
Amazon Polly has no Polly-specific consent clause at all. The only consent language we found is a general prohibition in the AWS Responsible AI Policy against depicting “a person’s voice or likeness without their consent” — an unnumbered bullet covering the whole AI service family.
Nothing published
MiniMax gates cloning on the developer’s verified identity, which is a real abuse control on the account. But we found no voice-owner consent requirement published anywhere in its developer documentation, across two independent verification passes. The cloning flow takes an uploaded audio file and returns an identifier.
What this means for you: if you are cloning your own voice, this section may not change your decision. If you are cloning anyone else’s — a client, a colleague, a voice actor — the provider that verifies consent protects you as much as the speaker, because the consent record exists independently of your word for it.
What cloning actually costs
| Provider | Entry route | Self-serve? | Sample required |
|---|---|---|---|
| Cartesia | $5/month Pro | Yes | 10 seconds (instant); 30 minutes (professional) |
| ElevenLabs | $6/month Starter | Yes | Instant from Starter; professional from $22 Creator |
| Speechify | $10/month API Starter | Yes | Speaker must record a phrase the vendor issues |
| MiniMax | Pay-as-you-go, no subscription | Yes | Not documented as a fixed requirement |
| Google Cloud | $60 per 1M characters — allow-listed | No | Audio plus a scripted recorded consent statement |
| Resemble AI | Contradictory: “first clone free, then $2 each” against a $1,000/month plan gate | Disputed | Not reconciled on any single page |
| Azure Speech | Not priceable without access | No | Recorded statement plus biometric match |
| OpenAI | No price published | No | Two separate recordings |
| Murf AI | Enterprise engagement | No | Studio-quality recordings, multi-week turnaround |
| Amazon Polly | No price published | No | Not published |
Four of these ten publish no price for cloning at all. Amazon, Azure and OpenAI route the question to sales without a figure; Resemble publishes two figures that cannot be reconciled — a changelog dated May 8, 2026 saying “first clone free, then $2 each”, and documentation stating “The Voice Cloning API requires a Business plan or higher”, that plan being $1,000 a month with bullets that never mention voice. We publish both and label the conflict rather than picking one.
This page ranks providers on consent and licensing. What the process actually requires of you — how much audio each tier needs, how long it takes, the exact consent script you must record, and what happens to the clone afterwards — is set out separately in our guide to how AI voice cloning works.
Instant against professional cloning
Where a provider offers both, the split is consistent: a short sample for a fast approximate clone, or a long recording session for a higher-fidelity one.
Cartesia documents the clearest version: instant cloning from a 10-second sample on the $5 Pro plan, professional cloning from 30 minutes of audio on the $49 Startup plan, with professional clones capped at 2 on Startup and 4 on Scale.
ElevenLabs splits by plan rather than sample: instant from $6 Starter, professional from $22 Creator, with professional restricted to your own voice.
We cannot tell you whether professional cloning sounds meaningfully better than instant cloning at any of these providers. That comparison requires listening, and our benchmark has not run. What the documentation establishes is the input cost — seconds against half an hour of recording — and the price step.
Feature matrix
| Provider | Third-party voices | Clone persists | Export the clone | Watermarked output |
|---|---|---|---|---|
| ElevenLabs | Instant only — professional is own-voice | Yes | No — clones cannot be downloaded | Not documented |
| Speechify | Yes, with verified consent | Yes | Consumer app: no MP3 download | Not documented |
| Cartesia | Yes, on your warranty | Yes — survives model migration | Not documented | Not documented |
| MiniMax | Yes | No — expires unless used in synthesis | Not documented | Mainland endpoint only |
| Resemble AI | Yes | Not documented as expiring | Not documented | Yes, on open-source output |
| Google Cloud | Yes, with recorded consent | Yes | Not documented | Not documented |
| Azure Speech | Yes, with biometric verification | Yes | Not documented | Not documented |
| OpenAI | Yes, two recordings required | Yes | Not documented | Anti-evasion duty published |
| Murf AI | Yes, with written consent | Yes | Not documented | Not documented |
| Amazon Polly | Brand Voice engagement | Yes | Not documented | Not documented |
What cloning is good and bad at today
Good at:
- Cloning your own voice cheaply. Two providers do it self-serve for under $7 a month, and one restricts professional cloning to exactly this case.
- Producing a defensible consent record, if you choose a provider that verifies rather than asserts.
- Surviving vendor model changes — Cartesia explicitly states clones carry across migrations with the same voice IDs.
Bad at:
- Being priced transparently. Four of ten publish no cloning price; a fifth publishes two that contradict each other.
- Portability. Where documented, clones generally cannot be exported — ElevenLabs states it outright. A cloned voice is a vendor lock-in.
- Persistence. MiniMax voices expire unless used in synthesis.
- Consistent safeguards. The gap between machine-verified consent and an unchecked warranty is enormous, and nothing in the pricing signals which you are buying.
Suitable and unsuitable use cases
| Use case | Fit | Which provider |
|---|---|---|
| Cloning your own voice for your own content | Strong | ElevenLabs $6/month or Cartesia $5/month |
| A product where your users clone their voices | Strong on consent, check the terms | Speechify verifies by recording — but its API terms appear to forbid end-user uploads its own docs design for |
| Cloning a client’s or actor’s voice commercially | Mixed | Use a provider that verifies consent; Murf’s written-consent regime suits contracted talent |
| Enterprise brand voice | Strong, at a price you must ask for | Amazon Brand Voice, Azure custom voice, or Murf — none publishes a figure |
| Cheapest possible cloning | Weak recommendation | MiniMax is cheapest and documents no voice-owner consent requirement. We do not recommend it on price alone |
| Cloning without vendor lock-in | Weak everywhere | Clones generally cannot be exported. Self-hosting an open-source model is the only route that avoids it |
If none of these fits
Resemble AI publishes MIT-licensed open-source speech models with an express commercial grant, self-hostable with no account and no usage caps. That is the only route here that avoids both the vendor lock-in of an unexportable clone and any per-clone fee — at the cost of running the models yourself, and with the consent obligation resting entirely on you.
The provider-by-provider detail behind every claim on this page is in the individual reviews. The ten-provider view across all dimensions is the flagship ranking, and the free-tier question is covered separately in best free AI voice generators — no free tier here includes cloning.
Frequently asked questions
Did you test how these cloned voices sound?
No. We created no accounts and cloned no voices. Every statement comes from vendor documentation read on September 1, 2026. Our listening benchmark is designed but not running, so nothing here compares clone fidelity.
What is the cheapest way to clone a voice?
Cartesia at $5/month is the cheapest self-serve subscription route, with ElevenLabs at $6. MiniMax’s pay-as-you-go rapid cloning is cheaper still and needs no subscription — but we found no voice-owner consent requirement published in its documentation, and we do not recommend cloning on price alone.
Can I clone someone else’s voice?
Legally that depends on your jurisdiction and their consent, which is not a question this page can answer. Technically, ElevenLabs forecloses it at the professional tier — own voice only. Several providers require a recorded consent statement from the speaker before a voice can be created. Others take your warranty and check nothing.
Which provider enforces consent most strictly?
Azure Speech, by documentation: a mandatory recorded statement for both tiers backed by biometric matching. Speechify has the strongest enforcement reachable self-serve, since Azure’s cloning is closed to anyone without a Microsoft account team.
Can I download or move a cloned voice?
Generally no. ElevenLabs states it directly: “You cannot download or export your voice clones as standalone files.” Treat a cloned voice as tied to the vendor you made it with.
Is any cloning available on a free tier?
No. No free tier among the ten providers we cover includes voice cloning.
Why is Resemble AI ranked fifth when it markets cloning heavily?
Because its documented access route cannot be reconciled: a changelog says the first clone is free and further clones cost $2 each, while its documentation states the cloning API requires a $1,000-a-month Business plan whose published bullets never mention voice. Both are official and neither references the other.
What about instant versus professional cloning?
The documented difference is the input: seconds of audio against roughly half an hour, at a higher price. Whether the result sounds meaningfully better is a listening question, and our benchmark has not run.
Sources and what we could not verify
Every claim on this page was read from official vendor sources on September 1, 2026, in a verification pass covering all ten providers with an adversarial re-check of every price and access claim. The per-provider detail, with full quotations, is in the individual reviews linked throughout. How we verify anything is described on the methodology page.
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- How any cloned voice sounds, or how closely it matches its source. No listening test has been run. This is the central question a cloning buyer has, and this page does not answer it.
- Cloning prices for four providers. Amazon Polly, Azure Speech and OpenAI publish none; Resemble publishes two that contradict each other.
- Whether MiniMax imposes any voice-owner consent requirement. Two independent passes searched its developer documentation and found none. Recorded as absent from the pages read.
- Brand Voice sample requirements, turnaround or minimum commitment. Amazon publishes none.
- Azure custom-voice pricing. Not reachable without the access registration.
- Whether Speechify API customers may let their own end users upload voice samples. Its API terms appear to forbid it while its documentation designs a flow around it.
- Whether cloned voices can be exported at most providers. Only ElevenLabs states the position explicitly, and its answer is no.
Change log
September 2, 2026 — First publication. All prices, access gates and consent requirements read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · This page shows no cloned-voice samples, because none exist: our benchmark has not run and we generate no audio. When it does, this page gains measured comparisons of clone fidelity and the ranking is revisited with a dated entry (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.