MiniMax vs Resemble AI 2026
Published September 2, 2026 · This is a documentation-based comparison. Every figure below was read from the two vendors’ own sources on September 1, 2026 and carries that check date. No listening test has been run, so nothing here compares how the two providers sound.
The short answer
One of these publishes what speech costs and the other does not publish it at all. MiniMax prices synthesis at $60 per million characters on its international platform, in real server-rendered tables. Resemble AI’s published rate card — Flex $0, Team $350, Business $1,000 a month — prices deepfake detection; text-to-speech, voice cloning and speech-to-speech appear nowhere on it (MiniMax pricing, Resemble pricing, both checked September 1, 2026).
Cloning access is the sharpest divide. MiniMax offers it on the standard API, though cloned voices expire unless used. Resemble’s cloning API is documented as requiring “a Business plan or higher” — the $1,000-a-month plan, whose own published bullets never mention voice at all.
Resemble’s real answer is free. It publishes MIT-licensed speech models with an express commercial grant, self-hostable with no account and no usage caps. For many buyers that, and not the hosted service, is the reason to consider it.
Comparison basis: official documentation only, checked September 1, 2026 — no accounts created, no API calls made, no audio generated, no scores assigned (how we verify).
| MiniMax | Resemble AI | |
|---|---|---|
| Published speech price | $60 per 1M characters (speech-2.8-turbo); $100 per 1M (speech-2.8-hd) | None. No per-character, per-second or per-minute speech rate is published anywhere |
| What the rate card prices | Speech synthesis, per model | Deepfake detection, intelligence, meetings, identity and watermarking |
| Cloning access | Standard API, identity-gated on the developer | Documented as requiring the $1,000/month Business plan |
| Cloning price | Charged on first use in synthesis | One figure exists, in a May 8, 2026 changelog: “first clone free, then $2 each” |
| Cloning lifecycle | Voices expire unless used in synthesis | Not documented as expiring |
| Output ownership | International: granted outright. Mainland: conditional on local law | Not stated. The single occurrence of “output” in the Terms is a restriction |
| Free route | No documented free speech tier | MIT-licensed open-source models, self-hosted, no account, no caps, express commercial grant |
| Watermarking | Mainland endpoint has a parameter; international has none | Applied by default to every open-source generation |
| Disclosure duty | Affirmative duty in the international terms | Not identified as an equivalent contractual duty |
| Deployment | Cloud API only | Cloud, plus documented on-premise and self-hosting |
| Audio quality | Not compared. Our listening benchmark has not run, so this page assigns no naturalness verdict to either provider. | |
On this page
Winners by category
Every winner below is a documentation-based judgment from published prices, terms and feature documentation. None is a test result, and the categories that require listening have no winner.
| Category | Winner | Why |
|---|---|---|
| Published pricing | MiniMax | Real per-model tables; Resemble publishes no speech rate at all |
| Free route | Resemble AI | MIT-licensed models with an express commercial grant beat having no free tier |
| Cloning on a small budget | MiniMax | Standard API against a documented $1,000/month gate |
| Cloning persistence | Resemble AI | MiniMax voices expire unless used in synthesis |
| Output ownership | MiniMax | Granted outright internationally; Resemble’s terms never assign it |
| Self-hosting | Resemble AI | Nothing MiniMax publishes comes close |
| Watermarking | Resemble AI | Applied by default on open-source output rather than left to the customer |
| Legal simplicity | Resemble AI | One entity and one governing regime against MiniMax’s two of each |
| Documentation coherence | Neither | MiniMax has undocumented model IDs and an undefined billing unit; Resemble markets a model its own docs list as end-of-life |
| Naturalness | No winner | Requires listening. Our benchmark has not run |
| Pronunciation and difficult text | No winner | Requires listening. Our benchmark has not run |
What each actually costs
MiniMax you can budget from published information. $60 per million characters on the turbo model, $100 on the HD model, in server-rendered tables, with a dedicated asynchronous endpoint for long text at the same rates. A 100,000-character chapter costs $6.00 on turbo. Two caveats travel with that: the mainland platform prices the same models in yuan per ten thousand characters, and the international audio subscription is denominated in “audio points”, a unit the documentation never defines.
Resemble you cannot budget from published information at all. Its pricing page carries a full self-serve rate card — Flex $0 a month, Team $350, Business $1,000, Enterprise custom — but the comparison rows on that page cover Resemble Detect, Resemble Intelligence, Resemble Meetings, Resemble Identity, Resemble Watermarker and platform access. No speech-synthesis product appears in any row. We checked the pricing page, the voice product pages, the documentation and the site’s own sitemap: no per-character, per-second or per-minute speech rate is published anywhere.
The one voice figure Resemble publishes sits in a dated changelog rather than the rate card. A post of May 8, 2026 states: “Voice cloning pricing changed so the first clone is free. Every user gets one clone before anything else, and only pays from the second one on, at $2 each.” We report it with its date and location and do not merge it into the rate card, because Resemble does not.
What this means for you: any hosted Resemble voice budget starts with a sales conversation. MiniMax’s starts with arithmetic.
Cloning: a standard API against a $1,000 gate
Resemble’s cloning gate is the single most consequential fact on this page for a small buyer. Its documentation states plainly: “The Voice Cloning API requires a Business plan or higher”, and for streaming, “This API is available to Business plans and above.” The Business plan is published at $1,000 a month, and its listed bullets — everything on Team, deployment configurability, calendar integration, security notifications, single sign-on, twenty seats — never mention voice, cloning or synthesis. The documented entry price for programmatic cloning is assembled from two documents that do not reference each other.
MiniMax offers cloning on the standard API, without an enterprise conversation. The catch is a timer rather than a price: “Voices produced by cloning and voice design are temporary: the fee is charged only on first use in speech synthesis”. A clone you create and do not immediately use in synthesis is a clone you may not have. If your workflow clones in a batch and synthesizes later, that gap is where voices disappear.
Neither enforces voice-owner consent in the product. MiniMax gates cloning on developer identity but publishes no voice-owner consent requirement anywhere in its developer documentation. Resemble markets consent tooling, but its Terms clause 2(a) reads — the typographical error is the vendor’s own — “Resemble may require consent form the individual or third party whose voice is being cloned. Consent needs to be verbal, unless otherwise stated by Resemble.” That is discretionary and sets a verbal standard, while its product page describes something stricter. Where a contract and a product page disagree, the contract governs.
Ownership and the open-source escape
MiniMax grants ownership; Resemble never mentions it. MiniMax’s international terms state: “As between you and us, and to the extent permitted by applicable laws, you retain your ownership rights in Client input and generated content” — though the same clause reserves a broad right to use that content to develop and improve its services, with no opt-out we could find.
Resemble’s hosted terms never affirm that you own the audio. We read the full Terms of Service: there is no clause assigning ownership of generated audio to the customer. A whole-word search returns exactly one occurrence of “output” in the entire contract, and it is a prohibition, in the acceptable-use clause, on using output “to train, improve, or otherwise further develop your own or any third party’s product, service, or deepfake detection model.” Clause 2(b) separately defines “Resemble AI Materials” broadly — including “audio, video … as well as all derivative works thereof” — as “owned by us”, granting only a “revocable, limited-purpose right to access and use”, and adding: “You are not permitted to download, copy or otherwise store any Resemble AI Materials.” We do not tell you what that adds up to, because the contract does not say and we do not fill silences.
The open-source path has none of this problem, and it is Resemble’s strongest offering. Its models are released under the MIT licence with an express commercial grant, quoted verbatim from the vendor’s own FAQ: “You can use them in commercial products, self-host, modify the weights, and ship to production — no royalties, no revenue share, no usage caps.” The quickstart is a pip install with “No API keys, no rate limits, no sign-ups.” We read that licence from the repository’s LICENSE file rather than a marketing badge — which matters, because a second Resemble model is badged “MIT OPEN SOURCE” on the site while its repository ships a community licence requiring companies over $10 million in revenue to buy a commercial licence.
Watermarking and disclosure
Both providers touch synthetic-media labelling, and they land on opposite sides of who does the work.
MiniMax puts the duty on you. Its international terms impose an affirmative obligation on the customer concerning the labelling of AI-generated content, in the user-obligations section of the agreement effective March 30, 2026. Set against that, we enumerated every property of the international text-to-audio request object and found no watermark parameter at all. The mainland endpoint does expose one. So on the platform most non-Chinese buyers will use, the obligation to label is yours and the tooling to do it in the audio is not provided by the API.
Resemble does it for you, on the free path. Its open-source repository states: “Every audio file generated by Chatterbox includes Resemble AI’s Perth (Perceptual Threshold) Watermarker.” That is a genuine and unusual commitment — the watermark is applied to free, self-hosted output where no commercial relationship compels it. Whether hosted API output carries the same watermark is not documented on any page we read, and we record that as unverified rather than assuming it carries across.
One further Resemble note for anyone evaluating on capability: the model its product page markets is listed as end-of-life on the hosted API in the same company’s documentation, which names a different model as current. Confirm which model your account actually runs before committing.
Voice quality, pronunciation and naturalness
We have not compared how these two providers sound, and we will not imply otherwise.
Our standardized listening benchmark is designed but not running. Until it does, this page carries no naturalness verdict, no pronunciation comparison and no scores for either provider. Both vendors publish quality claims of their own; those belong to them, and we do not repeat them as findings.
Choose MiniMax if · Choose Resemble AI if
Choose MiniMax if:
- You need a published price you can budget from. Resemble publishes no speech rate at all.
- You need cloning without a $1,000-a-month plan. MiniMax offers it on the standard API.
- You want ownership granted in the terms rather than left unaddressed.
- Long-form batch synthesis is the job, with a purpose-built asynchronous endpoint.
Choose Resemble AI if:
- You can self-host. MIT-licensed models with an express commercial grant and no caps are the best free route in our coverage.
- You need on-premise or air-gapped deployment. MiniMax offers nothing comparable.
- Watermarking by default matters to your organisation’s position on synthetic media.
- You want a single jurisdiction rather than two entities under two legal systems.
Sources and what we could not verify
Every figure on this page was read from the two vendors’ official sources on September 1, 2026. Full detail is in the individual reviews: MiniMax review and Resemble AI review. How we verify anything is described on the methodology page.
| Source | Used for |
|---|---|
| MiniMax international pricing | USD speech rates and units |
| MiniMax mainland pricing | CNY rates and the per-10,000-character unit |
| MiniMax API overview | Cloning lifecycle and identity gating |
| MiniMax text-to-audio reference | Request properties — establishing the absence of a watermark parameter |
| MiniMax international Terms | Ownership and use clause, disclosure duty (Effective March 30, 2026) |
| resemble.ai/pricing | The detection rate card and its comparison rows |
| Resemble cloning overview | The Business-plan requirement |
| Resemble Terms of Service | Clauses 2(a), 2(b) and 8 — consent, materials ownership, the output restriction |
| Chatterbox model page | MIT licence and the express commercial grant |
| Chatterbox repository | LICENSE file read directly; default watermarking statement |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- How either provider sounds. No listening test has been run.
- Any hosted Resemble speech rate. Not published on the pricing page, the product pages, the documentation or the sitemap. We publish no estimate.
- Whether Resemble’s changelog cloning price and its Business-plan API gate describe the same thing. Both are official; neither references the other.
- Whether hosted Resemble output carries the watermark. Documented for open-source generations only.
- Whether generated audio falls inside Resemble’s clause 2(b) definition of materials it owns. The contract neither includes nor excludes it expressly.
- What a MiniMax “audio point” is. Used to price its subscription and defined nowhere.
- Whether any MiniMax opt-out from model-development use exists. We found none documented on either platform.
- What
speech-01-hdandspeech-01-turbocost or do. Accepted by every MiniMax speech endpoint, documented nowhere.
Change log
September 2, 2026 — First publication as a documentation-based comparison. All figures read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · When our audio benchmark runs, this page gains a head-to-head listening test using the same prompts and documented conditions for both providers, and the naturalness and pronunciation rows gain real winners (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.