How AI Voice Cloning Works

✓ Docs-verified · Sep 2, 2026◌ Not audio-tested△ What we could not verify

Published September 2, 2026 · Every requirement quoted below was read from the vendor’s own documentation on September 2, 2026. This page covers what the process requires of you — audio, time, consent and what happens afterwards. It does not rank providers; that is the job of our voice cloning comparison. No listening test has been run, so nothing here says how closely any clone resembles anyone.

The short answer

Cloning a voice takes between 10 seconds and 30 minutes of audio, depending on which tier you use — and the two tiers are different products, not a quality slider. Instant cloning wants seconds; professional cloning wants half an hour or more.

Processing time spans five orders of magnitude. Azure’s personal voice states “Training time Less than 5 seconds”. Murf states “The training process typically takes 1 to 4 weeks.” Both describe creating a usable cloned voice.

The consent step is where these vendors differ most, and it is not a formality anywhere it exists. Speechify refuses the request if a different person reads the phrase. OpenAI fails it if you deviate from the script. Google will not let you substitute your own wording. MiniMax asks for nothing at all.

One provider will not let you clone another person’s voice even with their consent. Another answers “Can I clone a celebrity voice?” with “Yes” and an ethical caveat. Those are the two ends of this market.

Basis: vendor documentation only, checked September 2, 2026 — no accounts created, no voices cloned (how we verify).

On this page

What actually happens

Stripped of marketing, the documented process is the same four steps almost everywhere. What changes between vendors is how demanding each step is.

  1. You supply a sample. Between 10 seconds and 30 minutes of recorded speech, depending on the tier. Every vendor publishes format and size constraints alongside the duration, and several publish an upper limit as well as a lower one.
  2. You supply consent, or you assert it. This is the step that genuinely differs. Four providers make you record a fixed script and check it against your sample; two take a warranty in their terms and check nothing; one asks for nothing at all; the rest do not sell self-serve cloning.
  3. The vendor processes it. Anywhere from under five seconds to four weeks. Instant cloning is effectively immediate; professional cloning is a queued fine-tuning job.
  4. You get a voice identifier you call like any other voice — and, on some platforms, a retention clock you were not told about at the point of creation.

Two tiers, not one product. Most vendors selling cloning sell it twice: an instant tier built from seconds of audio, and a professional or fine-tuned tier built from half an hour or more. ElevenLabs names them Instant Voice Cloning and Professional Voice Cloning; Cartesia names them Instant Voice Clone and Pro Voice Clone; Azure names them personal voice and professional voice. The requirements below are per tier for that reason, because quoting one number for a vendor that sells two is how most comparisons get this wrong.

How much audio each provider needs

Published sample requirements, per tier, in each vendor’s own words — read September 2, 2026
ProviderTierSample required
ElevenLabsInstant“Record at least 1 minute of audio”; recommends “1–2 minutes of good audio”
ElevenLabsProfessional“The bare minimum we recommend is 30 minutes of audio”; recommends “30–180 minutes”
CartesiaInstant“as little as 10 seconds of audio”
CartesiaPro“30 minutes is the minimum”
Azure SpeechPersonal voice“a clean human voice sample between 5 - 90 seconds”
Azure SpeechProfessional voiceUtterances over 30 seconds are rejected: “Split the long audio into multiple files”
SpeechifyOne flow“Record 10-30 seconds of clean speech, under a minute and under 5MB, with no background noise”
MiniMaxOne flow“Duration: minimum 10 seconds, maximum 5 minutes”, up to 20 MB, mp3/m4a/wav
OpenAICustom voices“The audio samples must be 30 seconds or less”; at most 20 voices per organization
Google CloudInstant custom voiceReference audio capped at 10 seconds, recorded in the same environment as the consent audio
Resemble AIAPI flow“Audio needed: 10 seconds – 3 minutes”; upload path takes a “Single WAV file (≥10 seconds)”
Murf AISales-led“less than 90 minutes of high-quality, noise-free recordings” — a ceiling; no minimum published

Note the ones that publish a maximum as well as a minimum, because it is easy to assume more audio is always better. MiniMax caps the source at 5 minutes. Azure’s professional tier actively rejects individual utterances over 30 seconds and advises keeping them under 15. OpenAI caps samples at 30 seconds. Murf publishes only a ceiling and no floor at all.

MiniMax also wants a second, separate file. An optional example or prompt audio carries its own limit: “Duration: less than 8 seconds”. It is the only provider here documenting a two-file input.

Amazon Polly is absent from this table because its Brand Voice offering publishes no self-serve cloning mechanics at all. It is a sales-led engagement, and our Amazon Polly review covers what it does document.

How long it takes

This is the widest spread of any figure on this site, and it decides whether cloning fits your workflow.

Published processing time, verbatim — read September 2, 2026
Provider and tierStated time
Azure Speech — personal voice“Training time Less than 5 seconds”
Resemble AI — API flow“Training time: Under 1 minute”
Cartesia — Pro Voice Clone“Training then takes up to 3 hours.”
ElevenLabs — Professional“Generally fine-tuning takes 3-6 hours to complete, but it can sometimes take a bit longer, depending on the number of other PVCs queued for fine-tuning.”
Murf AI — sales-led“The training process typically takes 1 to 4 weeks.”

Under five seconds against four weeks is not a difference of degree. Azure’s personal voice and Resemble’s API flow are fast enough to sit inside a user-facing product; ElevenLabs and Cartesia’s professional tiers are overnight jobs; Murf’s is a project with a schedule. If you are planning a launch around cloned voices, this table is the one to plan against.

Two of these figures come with an honest caveat from the vendor. ElevenLabs notes its time depends on queue depth, and Murf’s own page frames its figure as a ceiling elsewhere, saying the clone is generated “within 4 weeks”. Neither vendor promises a number, and we are not turning their ranges into one.

Everything above is logistics. This is the part with consequences, and the vendors fall into three groups. Our voice cloning comparison ranks providers on where they land; this section describes what each one actually makes you do.

Speechify has the most thoroughly engineered consent gate in this market, and it publishes the engineering. Its documentation states plainly: “Consent is verified, not asserted.” The mechanism is that the vendor issues a phrase containing a per-request verification code, the speaker records it, and the recording is checked against the sample. Verbatim: “There is no create path that skips consent: consent_challenge_id and consent_recording are required fields, and the recording is retained as the consent record for the voice.” And the check has teeth: “The consenting speaker is the cloned speaker; a different voice reading the phrase is refused with consent_speaker_mismatch.”

OpenAI requires two recordings and matches them. A consent recording and a sample recording, with the voice required to match between them, and a script that cannot vary: “The consent audio recording must only include one of the following phrases. Any divergence from the script will lead to a failure.” One practical allowance is documented: “the consent can be used for multiple different voice creations if the same voice actor is making multiple attempts”.

Azure’s professional voice requires a recorded consent statement before fine-tuning, and Microsoft states the recording is also used to check the speaker is the same person as in the fine-tuning data. The script names both the speaker and the company, and the portal requires the submitted names to match what was spoken. Its personal voice tier separately requires explicit consent from every end user whose voice is cloned.

Google requires a consent recording capped at 10 seconds, recorded in the same environment as the reference audio, using a fixed script per language: “You can’t use a custom consent script instead of the default. You must use the provided consent statement script for your chosen language.”

Cartesia takes a policy warranty rather than a step in the flow: “You may only submit your own voice and audio recordings or those of others with explicit consent”. That sentence lives in its acceptable use policy, not in the cloning process. ElevenLabs’ instant tier asks you to tick a box: “Name and label your voice clone, confirm that you have the right and consent to clone the voice, then click”.

MiniMax documents no consent step whatsoever. The only gate on creating a clone in its cloning documentation is an account-level permission, surfaced as a status code: “2038: No cloning permission, please check account verification status”. No recorded consent statement, no identity check, no ownership warranty appears on its cloning pages. You upload audio and you get a voice.

The consent scripts, quoted

Where a fixed script exists, it is worth reading before you ask someone to record it, because it is a legal statement they are making about themselves.

Two things stand out. OpenAI and Google use the same sentence, differing only in the company name — two vendors independently converging on identical wording, which suggests it is doing legal work rather than product work. And Azure’s is the only one that is not an ownership claim. It asserts awareness and permission rather than ownership, and it names the company that will hold the result, which is a meaningfully different thing for the speaker to say.

Speechify’s is the only script that cannot be recorded in advance, because the verification code changes per request. That single design choice is what turns consent from a document into a check.

Whose voice you are allowed to clone

The two ends of this market are further apart than anything else on this page.

ElevenLabs’ professional tier refuses third-party cloning outright, consent or not. Verbatim: “No. You can only create a Professional Voice Clone of your own voice. Even with their consent, you cannot clone someone else’s voice. All Professional Voice Clones require a verification process to confirm that the voice belongs to you.”

Murf’s marketing FAQ answers the question “Can I clone a celebrity voice?” with “Yes”, followed by an ethical caveat rather than a prohibition: “However it can be deemed as unethical as the use of a celebrity voice without their permission can lead to intellectual property infringement or even identity theft.” No technical restriction is described anywhere on that page.

Those two positions come from the same market in the same week. If you are choosing a provider on governance rather than features, that contrast is the finding.

One structural note on Murf: its public API exposes no voice-cloning endpoint at all. The complete API reference index covers speech, streaming, voices, voice changer, translation, auth and dubbing — there is no create-clone operation. Cloning is a sales-led service whose only call to action is “Get in touch with our enterprise team to get started.”

What a clone actually is — a learned representation rather than a recording, and why the instant tier conditions a model at inference time while the professional tier changes its weights — is set out in our guide to how AI voice generators work.

What happens to the clone afterwards

The least-read part of the documentation, and the one most likely to surprise you.

MiniMax deletes your clone if you stop using it. Verbatim: “If a cloned voice is not used within 7 days, the system will delete it.” Nothing about that is unreasonable, and nothing about it is visible at the point of creation either. If you clone a voice for a seasonal campaign and come back three weeks later, it is gone.

Deletion on demand is documented by four providers, with a detail worth carrying in each case:

Speechify retains the consent recording deliberately, as the consent record for the voice. That is a reasonable design and it means a recording of the speaker exists on the platform for as long as the voice does.

Murf publishes where the data sits and not how long it stays. Its only data-handling statement on the cloning page is “Our AI models and voice data is stored in Amazon Web Services (AWS)”. No retention period and no deletion path appears there.

Where the documentation contradicts itself

Resemble AI’s ethics page and its terms disagree about whether consent is required. The ethics page states: “At the core of our ethical approach is a robust consent system that ensures voice cloning only occurs with the explicit consent of the individual.” The Terms of Service say: “Resemble may require consent form the individual or third party whose voice is being cloned” — the typo is in the original. “Ensures” and “may require” are not the same commitment, and the terms are the document that governs. Note separately that Resemble AI no longer sells voice generation.

Murf’s consent language is scoped to its own voice artists, not to your clones. Its ethics page states “Every Murf voice emerges from active collaboration and explicit consent from its creators” — a sentence that opens a section about artist payment and royalties. It describes how Murf’s stock voices were made. It does not describe a consent gate applied when a customer clones a third party, and it is easy to read as though it does.

Before you start

  1. Decide which tier you need first. Instant and professional cloning differ by three orders of magnitude in both audio and time. Choosing the tier settles most of the other questions.
  2. Record more audio than the minimum, within the maximum. Several vendors publish both, and exceeding the ceiling is a rejection rather than a warning.
  3. Read the consent script to the speaker before recording day. It is a statement about ownership and permission, and on three platforms a deviation fails the request outright.
  4. If you are cloning someone else, check whether you are allowed to at all — one provider prohibits it on its professional tier regardless of consent.
  5. Find the deletion path before you need it, and check whether there is a retention clock. One vendor deletes unused clones after seven days.
  6. Confirm the commercial position separately. Permission to clone is not permission to sell what you make with it — our guide to can you use AI voices commercially covers the licensing.
  7. Then choose a provider, using our voice cloning comparison, which ranks these ten on consent and licensing rather than on process.

Frequently asked questions

How much audio do I need to clone a voice?

Between 10 seconds and 30 minutes, depending on the tier. Instant cloning starts at 10 seconds with Cartesia and MiniMax and 10–30 seconds with Speechify. Professional cloning wants 30 minutes as a floor at both ElevenLabs and Cartesia, with ElevenLabs recommending 30–180 minutes.

How long does voice cloning take?

Azure states under 5 seconds for its personal voice and Resemble states under a minute for its API flow. Professional tiers are far slower: up to 3 hours at Cartesia, 3–6 hours at ElevenLabs depending on queue, and 1 to 4 weeks at Murf.

Can I clone someone else’s voice?

It depends entirely on the provider, and the range is extreme. ElevenLabs prohibits it on its professional tier even with consent. Murf’s FAQ answers the celebrity-cloning question with “Yes” plus an ethical caveat and no technical restriction. Four providers require a recorded consent statement that is checked against the sample.

Do I have to record a consent statement?

On four platforms, yes, using a fixed script you cannot alter — OpenAI, Google, Azure and Speechify. On Cartesia and ElevenLabs’ instant tier you warrant consent without recording anything. On MiniMax nothing is asked for at all.

Can I reuse one consent recording for several clones?

OpenAI documents that you can, if the same voice actor is making multiple attempts. Speechify’s cannot be reused, because the script contains a verification code issued per request.

Can I delete a cloned voice?

MiniMax, Resemble AI, Azure and Cartesia all document a deletion path. MiniMax additionally states the identifier cannot be reused afterwards, and deletes unused clones automatically after seven days. Murf publishes no retention period or deletion path on its cloning page.

Is instant cloning worse than professional cloning?

We cannot tell you that. It uses far less audio and far less processing, and every vendor positions the professional tier above the instant one — but those are documentation facts, not measurements, and no listening test has been run by this site.

Does cloning a voice mean I own it?

No, and the two questions are separate. Several providers never state that you own generated output at all. Our guide to can you use AI voices commercially sets out the governing clause for each.

Sources and what we could not verify

Every requirement, script and time quoted on this page was read from the vendor’s own documentation on September 2, 2026. Where a figure describes one tier of a multi-tier product, the tier is named in the row, because quoting a single number for a vendor that sells two is the most common error in this subject. Per-provider detail is in the ten reviews linked throughout, and the method is on the methodology page.

What we could not verify

Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 2, 2026:

Change log

September 2, 2026 — First publication. All requirements read from official vendor documentation on September 2, 2026 and dated accordingly.

Published September 2, 2026 · This page shows no audio and reports no listening results, because our benchmark has not run and does not cover cloning similarity. Cloning requirements and consent mechanisms change without notice; corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.