Resemble AI Review 2026
Published September 1, 2026 · Every fact below was read from Resemble AI’s own product pages, pricing page, developer documentation, Terms of Service, data processing addendum, trust centre, status page and public repositories on September 1, 2026, with the source linked beside the claim.
Quick verdict
Resemble AI presents itself as a deepfake-detection company that also sells voice, and its documentation follows that priority. The homepage title reads “Multimodal Deepfake Detection and Watermarking for Enterprise” and its heading is “Deepfakes are everywhere. So are we.” The Products navigation contains no speech-synthesis entry at all (resemble.ai, checked September 1, 2026). That is an observation about how the site is arranged today, not a claim about the company’s plans.
The consequence for a voice buyer is concrete: the published rate card prices detection products only. Voice cloning, text-to-speech, speech-to-speech and audio editing appear nowhere on it. Meanwhile the developer documentation states that “The Voice Cloning API requires a Business plan or higher” — a plan published at $1,000/month whose own feature list mentions no voice entitlement (cloning docs; pricing, both checked September 1, 2026). You cannot price a voice integration from the pricing page.
And the strongest route into this product is the free one. Resemble publishes open-source speech models under the MIT licence with an explicit commercial grant — no plan, no API key, no usage cap. For many buyers that is the actual recommendation on this page.
Verdict basis: official documentation only, checked September 1, 2026 — no account, no API call, no listening test, no numeric score (how we verify).
| Best for | Teams who want self-hosted, MIT-licensed speech synthesis with no vendor account, or enterprises buying deepfake detection and watermarking |
|---|---|
| Not best for | Anyone who needs to budget a hosted voice integration from published prices, or who needs a contract that states who owns the generated audio |
| Published pricing | Detection products only: Flex $0/mo, Team $350/mo, Business $1,000/mo, Enterprise “Custom”. No per-character, per-second or per-minute speech rate is published anywhere (pricing) |
| Voice cloning cost | One figure exists, in a dated changelog rather than the rate card: “first clone free, then $2 each” (May 8, 2026). The cloning API itself is gated to the $1,000/month Business plan |
| Output ownership | Not stated. The Terms of Service contain no clause assigning ownership of generated audio; the single occurrence of “output” in the contract is a restriction |
| Open-source route | Chatterbox and related models under the MIT licence, with an express commercial grant and no usage caps |
| Test status | Documentation-verified. Our listening benchmark has not run, so there is no independent voice-quality score and no measured latency on this page. |
Every price, licensing term and feature claim below was read from Resemble AI’s official properties on September 1, 2026, with the source linked beside the fact. Our standardized audio benchmark is designed but not yet running, so this review contains no listening-test claims, no naturalness verdicts and no scores. Where Resemble makes a quality, latency or accuracy claim, it is quoted and attributed to Resemble.
Strongest advantages:
- A genuinely free, genuinely commercial open-source path: the models are MIT-licensed with an express grant to “use them in commercial products, self-host, modify the weights, and ship to production — no royalties, no revenue share, no usage caps” (Chatterbox page, checked September 1, 2026).
- Self-hosting with no account at all: “No API keys, no rate limits, no sign-ups”, installed from pip.
- Audio watermarking is applied by default on open-source output rather than sold as an upsell — unusual, and aligned with the company’s stated position on synthetic media.
- Deployment options extend to on-premise, which very few providers in this comparison document at all.
- A public status page and a public trust centre, both of which disclose more than the marketing pages claim.
Important disadvantages:
- No hosted speech price is published: no per-character, per-second or per-minute rate exists on any official page.
- The Terms of Service never state who owns generated audio, and the one clause mentioning output is a restriction on using it.
- The voice-cloning API is gated to a $1,000/month plan whose published feature list contains no voice entitlement.
- The model the product page markets is listed as end-of-life on the hosted API in the same company’s documentation.
- A model badged “MIT OPEN SOURCE” on the site ships a repository licence that is not MIT and that requires larger companies to buy a commercial licence.
- Consent is marketed as verifiable and written as discretionary and verbal.
- The homepage compliance badge and the trust centre describe different SOC 2 states.
On this page
What kind of company this is now
Resemble AI began as a voice-synthesis company and is well known as one. That is not how its website presents it today, and a buyer evaluating it for voice work should know that before reading anything else.
As checked on September 1, 2026, the homepage document title is “Multimodal Deepfake Detection and Watermarking for Enterprise | Resemble AI” and its top-level heading is “Deepfakes are everywhere. So are we.” The Products navigation lists deepfake detection, intelligence, watermarking, identity registration, meetings protection, a Chrome extension and integrations. No speech-synthesis product appears in it.
The voice products still exist and still have documentation. They simply are not what the front door sells. Every finding below follows from that split — the pricing, the plan gating, and the gap between what the product pages market and what the API documentation says is running.
Who should use it — and who should look elsewhere
Choose Resemble AI if:
- You want to self-host speech synthesis under a permissive licence. The open-source models are MIT-licensed with an explicit commercial grant, installable from pip, with “no API keys, no rate limits, no sign-ups” (Chatterbox, checked September 1, 2026). Nothing else in our coverage offers that.
- Watermarked output is a requirement rather than an annoyance. Open-source generations carry Resemble’s perceptual watermark by default, which is a defensible position if your organisation has committed to labelling synthetic audio.
- You are buying detection, not just synthesis. The published rate card is built for exactly that, from a $0 Flex tier upward, and it is the part of the business the documentation serves best.
- You need on-premise deployment. Resemble documents deployment options extending to on-premise, which most providers here do not offer at any price.
Look elsewhere if:
- You need to budget a hosted voice integration. No speech rate is published. If you need a number before you can get approval, Amazon Polly publishes $4.00 per 1 million characters (pricing, checked September 1, 2026).
- Your legal review needs the contract to say you own the audio. It does not say so. Amazon Polly states it in Service Terms 50.2; several others state it plainly too.
- You want self-serve cloning at a consumer price. The cloning API is documented as requiring a Business plan or higher, published at $1,000/month. ElevenLabs documents cloning from $6/month (pricing, checked September 1, 2026).
- You need a stated language list you can rely on. Official Resemble pages give several different language counts, and no page enumerates the largest one — see languages.
Current price, free plan, commercial rights
The short version: the published rate card prices deepfake detection. Voice has one published price, in a changelog entry, and an API gate to a $1,000/month plan. The genuinely free route is the open-source one, and it is the only path with an express commercial grant.
| Plan | Monthly | Annual | What the plan rows actually cover |
|---|---|---|---|
| Flex | $0 /mo | — | The comparison rows on this page are Resemble Detect, Resemble Intelligence, Resemble Meetings, Resemble Identity, Resemble Watermarker and platform/API access. No speech-synthesis product appears in any row. |
| Team | $350 /mo | $280 /mo billed annually | |
| Business | $1,000 /mo | $800 /mo billed annually | |
| Enterprise | Custom | Talk to Sales |
Quoted from resemble.ai/pricing, checked September 1, 2026. The annual rows carry the vendor’s own savings wording (“save $840/yr” on Team, “save $2,400/yr” on Business). One usage rate is published on that page, for detection rather than synthesis: Resemble Watermarker decode at “$0.00020 per call”.
The one published voice price is in a changelog
A voice-cloning price does exist on the official site, but not on the rate card. A dated changelog post of May 8, 2026, titled “Voice cloning pricing: first clone free, then $2 each”, states verbatim: “Voice cloning pricing changed so the first clone is free. Every user gets one clone before anything else, and only pays from the second one on, at $2 each.” It adds: “Card capture is deferred until after that first free clone, so you can try cloning without entering payment details up front” (changelog, May 8, 2026, checked September 1, 2026).
We report that with its date and location and do not merge it into the rate card, because Resemble does not. Note also that it sits beside a documented plan gate that points the other way, described next.
The voice API is gated to a plan that does not mention voice
The developer documentation is explicit: “The Voice Cloning API requires a Business plan or higher”, and for streaming, “This API is available to Business plans and above” with the note “The WebSocket API is limited to Business plan customers” (cloning overview, streaming docs, checked September 1, 2026).
The Business plan is published at $1,000/month. Its listed bullets are “Everything on Team plan”, deployment configurability, org-wide calendar integration for Meetings, Meetings security notifications, Single Sign-On and 20 team seats. None of them mentions voice, cloning or synthesis.
What this means for you: the documented entry price for programmatic voice cloning at Resemble is $1,000/month, and you would not learn that from the pricing page. That is the single most important budgeting fact on this review, and it is assembled from two documents that do not reference each other.
Commercial rights differ sharply between the two routes
For the hosted service, no clause grants commercial use of generated audio and none prohibits it. The Terms of Service restrict sublicensing and “commercial time-sharing or service bureau use”, and forbid using output to build a competing product — but they contain no commercial-use grant of the ordinary kind.
For the open-source models, the grant is explicit and generous. From the Chatterbox FAQ, verbatim: “Yes. Chatterbox, Chatterbox Multilingual, and Chatterbox Turbo are released under the MIT license. You can use them in commercial products, self-host, modify the weights, and ship to production — no royalties, no revenue share, no usage caps.”
What this review covers
This review covers Resemble AI’s speech-synthesis offering as documented today: the hosted API and its plan gates, the published rate card and what it does and does not price, the open-source models and their licences, the watermarking behaviour, the consent requirements, and the contract terms governing output, data and jurisdiction. It also covers the detection products where they bear on what a voice buyer must purchase.
It does not evaluate how any Resemble voice sounds. We created no account, made no API call and generated no audio. Rows a review of ours normally carries once our benchmark runs — Tested, Model tested, Voice tested, Benchmark cost — are absent here rather than filled with invented values.
Document dates matter on this page and are given wherever quoted: the Terms of Service carry “Last Update: July 30, 2024” with “Last Review: January 1, 2026”, and the data processing addendum is dated March 17, 2026.
Benchmark audio status
Samples are pending. Our standardized audio benchmark is designed but not running, so this page publishes no audio and no empty player (methodology). What we verified instead is documentation: plan contents and gates, model status on the hosted API, repository licences read from the LICENSE files themselves, consent and watermarking wording, retention terms in the data processing addendum, and the published status and trust pages, all read on September 1, 2026.
Resemble publishes accuracy and adoption figures for its detection products. We did not verify those independently and report no number from them.
Assessment by dimension
No numeric scores appear below. The judgments are qualitative and drawn from what Resemble documents; vendor claims are labelled as such.
| Dimension | What the documentation supports |
|---|---|
| Open-source offering | Strong, and the best in this comparison. MIT licence, express commercial grant, no caps, no account, on-prem capable. |
| Hosted pricing transparency | Weak. No speech rate published anywhere; the documented cloning gate sits on a $1,000/month plan that never mentions voice. |
| Output ownership | Not stated. No assignment clause exists; the only occurrence of “output” in the contract is a restriction. |
| Consent tooling | Contested. Marketed as explicit and verifiable; written in the contract as discretionary and verbal. |
| Watermarking | Strong on open source. Applied to every generation by default. Hosted behaviour is undocumented. |
| Model clarity | Weak. The model the product page markets is marked end-of-life on the hosted API in the same company’s docs. |
| Language coverage | Unresolvable from official pages. Several different counts; the largest is enumerated nowhere. |
| Deployment options | Strong. Cloud, private and on-premise are documented, which is rare here. |
| Voice quality | Not assessed. We have run no listening test. |
Models: what the site markets and what the API runs
These are not the same thing, and the discrepancy is documented by Resemble itself.
The developer documentation states: “Resemble Ultra is the current text-to-speech model for new and upgraded voices”, and “All previous Resemble text-to-speech models have reached end of life. Voices that still use one of these versions can no longer generate audio and must be upgraded to Resemble Ultra.” Its end-of-life table lists Chatterbox (tts-v4), Chatterbox-Turbo (tts-v4-turbo) and Chatterbox Multilingual (tts-v4) as end of life (model versions, checked September 1, 2026).
The text-to-speech product page, meanwhile, still leads with Chatterbox: “Chatterbox delivers real-time speech synthesis, custom pronunciation, and PerTh watermarking” (text-to-speech product page, checked the same day). The status page monitors the sole text-to-speech component as “Resemble Ultra (HTTP)”.
Resemble Ultra has no product or model page on the public site. The only description we found is a changelog entry of April 29, 2026 announcing that “Resemble Ultra (powered by xAI) is live”.
What this means for you: if you evaluate Chatterbox’s published characteristics and then buy the hosted API expecting them, you may not be buying that model. Confirm which model your account runs before committing, and note that the open-source Chatterbox you can self-host is a different proposition from the hosted service that lists it as retired.
The open-source route
This is the part of Resemble’s offering we would point most readers at, and it is documented plainly.
The company states its position directly: “We build in the open. You use it for free.” and “PerTh, Resemblyzer, Chatterbox are free and open source. No usage limits, no rate caps — run them wherever you want, including on-prem” (start here, checked September 1, 2026). The quickstart is a pip install with “No API keys, no rate limits, no sign-ups.”
We read the licence from the repository rather than trusting the marketing badge: the Chatterbox LICENSE file is an MIT Licence, “Copyright (c) 2025 Resemble AI”, and GitHub reports its SPDX identifier as MIT (repository, checked September 1, 2026, 26,226 stars at that time).
One badge does not match its licence file
Resemble’s models index badges DramaBox as “MIT OPEN SOURCE”, and its DramaBox page describes it as “Open source and ready for its close-up” while naming no licence and linking no repository.
The repository’s LICENSE file is not an MIT licence. It is titled “LTX-2 Community License Agreement” with a licence date of January 5, 2026 and the Licensor named as Lightricks Ltd. GitHub reports the SPDX identifier as NOASSERTION and Hugging Face reports “license: other”. Clause 2 of that licence states that entities “with annual revenues of at least $10,000,000 (the ‘Commercial Entities’) are required to obtain a paid commercial use license”.
This is not a theoretical concern: a changelog entry of June 10, 2026 records that Voice Design moved to the DramaBox model. We report the badge and the licence file as we found them, both linked, and draw no conclusion about how the discrepancy arose. If your company clears $10 million in revenue and you are relying on a DramaBox-derived path, read that licence yourself before shipping.
Voice quality and naturalness
We will not tell you how Resemble AI sounds. No listening test has been run for this review.
Resemble publishes quality and accuracy claims, particularly for its detection products, and markets its synthesis models on realism. Those are vendor claims and we do not repeat them as findings. What documentation settles is the deployment envelope — formats, watermarking, model versions and plan gates — which is what this review actually delivers.
Voice cloning and consent
Resemble markets consent tooling as a differentiator, and it is the area where marketing and contract diverge most sharply.
The Terms of Service, clause 2(a), state verbatim — the typographical error is the vendor’s own: “Resemble may require consent form the individual or third party whose voice is being cloned. Consent needs to be verbal, unless otherwise stated by Resemble.” That is discretionary (“may require”) and sets a verbal standard.
The voice-creation product page describes something stricter, requiring explicit verifiable consent for a professional clone. Both statements are official and current, and Resemble does not reconcile them. Where a contract and a product page disagree, the contract is what governs.
Clause 2(a) also allocates the legal risk clearly, and this sentence is worth reading before any cloning project: “You represent and warrant that you own your Content or have the necessary licenses, rights, consents and permissions to grant the license set forth herein and that its provision to Resemble AI or Resemble AI’s use thereof will not violate applicable laws, the copyrights, privacy rights, publicity rights, trademark rights, contract rights or any other intellectual property rights or other rights of any person or entity.” The same clause adds: “You agree to pay all royalties, fees and any other monies owing to any person by reason of the Content uploaded, displayed or otherwise provided by you.”
On price and access, see the changelog figure and the Business-plan gate, which point in different directions.
Watermarking
Every open-source generation is watermarked by default. The Chatterbox repository states verbatim: “Every audio file generated by Chatterbox includes Resemble AI’s Perth (Perceptual Threshold) Watermarker” (repository, checked September 1, 2026).
That is a genuine and unusual commitment: the watermark is applied to free, self-hosted output where no commercial relationship compels it. Whether the same behaviour applies to hosted API output is not documented on any page we read, and we record that as unverified rather than assuming it carries across.
The watermarker is also sold as a product in its own right, with a published decode rate of “$0.00020 per call” on the pricing page.
Languages and accents
Official Resemble pages give several different language figures, and we could not settle them.
The text-to-speech product page cites 100 languages for the managed platform, in three places, with no enumeration published anywhere we could find. The FAQ on that same page gives 23 with a named list. The open-source repository README gives 23 with a different named list.
We publish the discrepancy rather than a number of our own choosing. If a specific locale decides your purchase, verify it directly against the model documentation for the exact model your account runs — which, per models, may not be the one the product page advertises.
API, deployment, compliance
Deployment options are a real strength here: Resemble documents cloud, private-cloud and on-premise environments, and advertises a “99.9% uptime SLA” for the cloud service (integrations and environments, checked September 1, 2026).
Two disclosures on Resemble’s own operational pages sit below what the marketing pages claim, and both are to the company’s credit for being published at all.
SOC 2. The homepage carries a badge reading “SOC 2 Type II — Independently audited security controls covering availability, confidentiality, and data integrity.” The trust centre states: “Resemble AI is currently in our SOC 2 Type 2 observation period” (trust centre, checked September 1, 2026). Both were live on the same day. An observation period is not a completed Type II audit, and a buyer whose procurement depends on that distinction should ask for the report directly.
Uptime. Against the advertised 99.9% SLA, the status page on September 1, 2026 showed a 90-day figure of 99.425% for the Safety & Detection API, with Deepfake Detection at 99.805% and Identities at 99.044%, while web portals showed 100% and 99.999% (status page).
Generation speed and latency
We verified no vendor latency figure for the speech products in a form we would publish, and we have run no measurement of our own. Resemble markets real-time synthesis; we do not convert a marketing adjective into a number.
This section will carry measured figures when our benchmark runs.
Pricing and normalized cost
What you pay, and what you get
There is nothing to normalize for hosted speech, because no speech rate is published. That is itself the finding. The plan prices above buy detection products and platform access; the documented route to programmatic voice cloning runs through the $1,000/month Business plan.
The important limitation: you cannot build a budget from the pricing page
For every other provider in our coverage, a buyer can compute an approximate monthly cost from published figures before contacting anyone. Here they cannot. The rate card prices a different product line; the one published voice figure lives in a changelog; and the API gate is documented somewhere else again. Any hosted voice budget at Resemble starts with a sales conversation.
Who gets the best value
- Self-hosting teams: excellent value. MIT-licensed models with a commercial grant and no caps cost nothing but infrastructure and engineering time.
- Enterprises buying detection plus voice: potentially good value, because the Business plan you would need for the cloning API also carries the detection entitlements — but you must price it with sales.
- Small teams wanting hosted voice: poor value. A documented $1,000/month gate for cloning is far above the market entry point.
Commercial usage, licensing, privacy
Who owns what you generate: the contract does not say
We read the full Terms of Service. There is no clause assigning ownership of generated audio to the customer. A whole-word search of the complete contract text returns exactly one occurrence of “output”, and it appears in clause 8, “Acceptable Use and Conduct”, as a prohibition:
“accesses the Services or otherwise uses any output generated from the Services (including without limitation audio files) to train, improve, or otherwise further develop your own or any third party’s product, service, or deepfake detection model.”
Clause 2(a) defines “Content” as what you supply — input audio and related data — not what the service returns. Clause 2(b) defines “Resemble AI Materials” broadly, as “all other materials and services provided by or through Resemble AI … including, but not limited to, the API, software, all informational text, software documentation, design of and ‘look and feel’, layout, photographs, graphics, audio, video, messages, design and functions, files, documents, images, or other materials, as well as all derivative works thereof”, and states these “are owned by us or our licensors or service providers”. It grants only a “non-transferable, non-sublicensable … non-exclusive, revocable, limited-purpose right to access and use” those materials, and adds: “You are not permitted to download, copy or otherwise store any Resemble AI Materials.”
We are not going to tell you what that adds up to, because the contract does not say and we do not fill silences. What we will say is the practical consequence: Resemble’s hosted terms never affirm that you own the audio you generate, and that is a materially weaker position than several competitors offer in a single sentence. If hosted Resemble output is going into commercial work, get written clarification before you rely on it. The open-source path does not have this problem — the MIT licence and its express commercial grant settle it.
Data, retention and training
The data processing addendum, dated March 17, 2026, documents retention defaults that keep voice data. Its Annex 1 records that voice cloning and model training data is “Retained for archival retrieval; deleted within 30 days of Controller’s deletion request”, and for text-to-speech inference that “Input text and generated audio retained for archival retrieval; deleted within 30 days of Controller’s deletion request”.
In other words, deletion is request-driven rather than automatic, on a 30-day clock from the request. Plan for that if you process sensitive scripts.
Jurisdiction
The Terms of Service, clause 16(a), select “the laws of the Province of Ontario, Canada”, with arbitration “under the rules of The ADR Institute of Canada (ADRIC)” and “The place of arbitration will be Toronto, Ontario, Canada.” The data processing addendum selects a different governing law in its own clause 15.4. Which document governs a given dispute is a question for your counsel, not for us — but the split is worth knowing before signing.
Pros
- The best open-source offering in our coverage: MIT-licensed models with an express commercial grant, no royalties, no usage caps.
- Self-hosting with no account, no API key and no rate limit, installable from pip.
- On-premise and private-cloud deployment documented, which almost nothing else here offers.
- Audio watermarking applied by default to open-source output rather than sold as an add-on.
- A public status page and trust centre that disclose real operational figures.
- A published, self-serve rate card for the detection products, starting at $0.
- Clear risk allocation in the cloning clause, so you know exactly what you are warranting.
Cons
- No published price for hosted speech synthesis of any kind.
- The Terms of Service never state who owns generated audio, and the sole mention of output is a restriction.
- Programmatic voice cloning is documented as requiring a $1,000/month plan whose feature list never mentions voice.
- The model marketed on the product page is listed as end-of-life on the hosted API in the same company’s documentation.
- A model badged “MIT OPEN SOURCE” ships a non-MIT licence requiring larger companies to buy a commercial licence.
- Consent is marketed as explicit and verifiable but written as discretionary and verbal.
- The homepage SOC 2 badge and the trust centre describe different audit states.
- Measured 90-day uptime on the detection API sat below the advertised 99.9% SLA on the day we checked.
- Language counts differ across official pages, and the largest is enumerated nowhere.
- Retention of input text and generated audio is request-driven deletion on a 30-day clock, not automatic.
Best use cases
| Use case | Fit | Why |
|---|---|---|
| Self-hosted synthesis | Strong | MIT licence, express commercial grant, no caps, runs on your own infrastructure |
| Air-gapped or on-prem deployment | Strong | On-premise is documented; the open-source path needs no vendor connection at all |
| Synthetic-media governance programmes | Strong | Watermarking and detection are the company’s current focus and its best-documented products |
| Enterprise voice with a procurement team | Mixed | Capable, but every number comes from sales and the ownership question needs answering in writing |
| Small-team hosted voice work | Weak | A documented $1,000/month gate for the cloning API is far above the market entry point |
| Anything needing a fixed published rate | Weak | None is published for speech |
And the inverse, because a recommendation without one is not worth much: if you want a hosted voice product you can price, buy and own the output of, without a sales call or a licence review, this is the wrong provider on today’s documentation — and the same company’s open-source release is the better answer.
Alternatives
Organized by the reason you would leave.
- You need a published price you can budget from. Amazon Polly lists $4.00 per 1 million characters with no subscription (pricing, checked September 1, 2026).
- You need ownership stated in the contract. Amazon Polly states it in Service Terms 50.2.
- You need self-serve cloning at a consumer price. ElevenLabs documents cloning from $6/month (pricing, checked September 1, 2026).
- You want an editor rather than an API. Murf AI sells a studio workflow from $29/month with a quotable commercial grant (pricing, checked September 1, 2026).
- You want the whole market in one view. The ranking of ten providers compares pricing, free tiers, licensing and cloning side by side.
Direct comparison links
Head-to-head pages pairing Resemble AI against individual competitors are on this site’s roadmap but not published yet, and we do not link to pages that do not exist. Until they are live, the flagship ranking carries side-by-side pricing, licensing and feature tables for all ten providers we cover.
Final verdict
Choose Resemble AI when you want to self-host. The open-source release is the strongest thing in this review by a distance: permissively licensed, commercially granted in writing, watermarked by default, and free of accounts, keys and caps. For a team with the engineering capacity to run models themselves, it is a genuinely good answer, and it is unusual for a commercial vendor to publish one.
Do not choose the hosted service when you need to know what it costs, or who owns what it produces. Neither question is answered by the published documentation: the rate card prices a different product line, the one voice figure lives in a changelog, the cloning API is gated to a $1,000/month plan that does not mention voice, and the contract never affirms that generated audio is yours. Those are not gaps we are inferring — they are what a complete read of the official pages produces.
What would change this verdict is a page. A published speech rate and a one-sentence ownership clause would move Resemble’s hosted offering substantially up this comparison without changing a line of its technology. Separately, when our audio benchmark runs, this page gains measured comparisons and the assessment table gains scores that mean something. Until then this is a verdict about documentation, and the documentation is the problem.
Frequently asked questions
Did you actually test Resemble AI’s voices?
No. We created no account, made no API call and generated no audio for this review. Every statement here comes from Resemble AI’s own documentation, repositories and status pages, read on September 1, 2026. Our listening benchmark is designed but not running (methodology).
How much does Resemble AI cost for text-to-speech?
No price is published. The rate card on the pricing page covers detection products: Flex $0/mo, Team $350/mo, Business $1,000/mo, Enterprise custom. No per-character, per-second or per-minute speech rate appears on any official page we read, and we will not estimate one.
What does voice cloning cost?
Two official answers point different ways. A changelog entry of May 8, 2026 says “first clone free, then $2 each”. The developer documentation says “The Voice Cloning API requires a Business plan or higher” — that plan is published at $1,000/month.
Do I own the audio I generate?
The hosted Terms of Service do not say. There is no ownership assignment clause, and the only occurrence of “output” in the contract is a restriction on using it to train competing products. The open-source models are different: MIT licence, with an express commercial grant.
Is Resemble AI free?
The open-source models are genuinely free and commercially usable, self-hosted, with no account and no caps. The hosted platform has a $0 Flex plan, but its published rows cover detection products rather than speech synthesis.
Is the audio watermarked?
Every open-source generation is, by the repository’s own statement. Whether hosted API output carries the same watermark is not documented on any page we read, so we make no claim about it.
What consent does voice cloning require?
The contract and the marketing differ. Terms of Service clause 2(a) says Resemble “may require consent” and that “Consent needs to be verbal, unless otherwise stated by Resemble”. The product page describes explicit verifiable consent for professional clones. The contract governs.
Is Resemble AI SOC 2 certified?
Its homepage badge says SOC 2 Type II with independently audited controls; its trust centre says the company is “currently in our SOC 2 Type 2 observation period”. Both were live on September 1, 2026. Ask for the report directly if this decides your purchase.
Sources and what we could not verify
Every changing fact on this page was read from official Resemble AI properties on September 1, 2026. Licences were read from the LICENSE files in the repositories themselves rather than from badges on marketing pages. Where two official pages disagree, both readings are published above and neither is averaged or silently reconciled.
| Source | Used for |
|---|---|
| resemble.ai | Homepage title, heading, product navigation, SOC 2 badge |
| resemble.ai/pricing | Plan names and prices, annual savings wording, comparison rows, watermarker decode rate |
| Changelog, May 8, 2026 | The one published voice-cloning price and the deferred card capture |
| Cloning overview docs | The Business-plan requirement for the Voice Cloning API |
| Streaming docs | The Business-plan limitation on the WebSocket API |
| Model versions docs | Resemble Ultra as current model; end-of-life table for the Chatterbox identifiers |
| Text-to-speech product page | Chatterbox marketing wording, language claims, open-source deployment note |
| Terms of Service | Clauses 2(a), 2(b), 2(c), 2(d), 8, 16(a) — content licence, materials ownership, consent, output restriction, jurisdiction (Last Update July 30, 2024; Last Review January 1, 2026) |
| Chatterbox model page | MIT licence FAQ and the express commercial grant |
| Chatterbox repository | LICENSE file read directly; default watermarking statement |
| Models index | The DramaBox “MIT OPEN SOURCE” badge |
| Start here | Open-source positioning and the no-caps statement |
| Integrations and environments | Deployment options and the advertised uptime SLA |
| Trust centre | SOC 2 observation-period statement |
| Status page | 90-day uptime figures and the monitored text-to-speech component name |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026, recorded rather than guessed:
- Any hosted rate for text-to-speech, speech-to-speech or audio editing. Checked the pricing page, the voice product pages, the documentation and the site’s own sitemap. No such rate is published. We publish no estimate.
- Whether the changelog cloning price and the Business-plan API gate describe the same thing. Both are official; neither references the other.
- Whether hosted API output carries the PerTh watermark. Documented for open-source generations only.
- Which language count is correct. Official pages give several, and the largest is enumerated nowhere.
- What Resemble Ultra actually is. It has no product or model page; the only description is a one-line changelog entry.
- How the DramaBox badge and its repository licence came to differ. We report both as found and draw no conclusion.
- The current SOC 2 state. Two official pages describe it differently on the same day.
- Whether generated audio falls inside the clause 2(b) definition of materials owned by Resemble. The contract neither includes nor excludes it expressly. We state the silence rather than fill it.
- Vendor latency figures for the speech products in a form we would publish.
- Resemble’s published detection accuracy and adoption figures. Not independently verified; no number from them is reported here.
Change log
September 1, 2026 — First publication. All prices, terms, model statuses, licences and operational figures read from official Resemble AI sources on this date; repository licences read from the LICENSE files directly.
Published September 1, 2026 · Benchmark audio, measured costs and scores will be added with a dated entry here when the audio benchmark runs (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.