Cartesia vs MiniMax 2026
Published September 2, 2026 · This is a documentation-based comparison. Every figure below was read from the two vendors’ own sources on September 1, 2026 and carries that check date. No listening test has been run, so nothing here compares how the two providers sound.
The short answer
These are the two cheapest serious speech APIs in our coverage, and both ask for something unusual in return. Cartesia sells a commercial licence and instant cloning for $5 a month. MiniMax prices synthesis at $60 per million characters on its international platform — low for long-form work, and published openly in real tables (Cartesia pricing, MiniMax pricing, both checked September 1, 2026).
Both take rights over your content by default, and neither makes it easy to decline. Cartesia’s licence over inputs and outputs is irrevocable, perpetual and sublicensable, with a manual, forward-only opt-out form. MiniMax reserves a broad use right and we found no opt-out documented at all on either of its platforms.
The complexity is different in kind. Cartesia’s is temporal — a deprecation calendar with an alias-level sunset seven weeks out. MiniMax’s is jurisdictional — two platforms under two entities, two governing laws, two currencies and two billing units.
Comparison basis: official documentation only, checked September 1, 2026 — no accounts created, no API calls made, no audio generated, no scores assigned (how we verify).
| Cartesia | MiniMax | |
|---|---|---|
| Entry price | $0 Free (no commercial licence, no cloning); $5/month Pro adds both | No plan. Pay-as-you-go, with no documented free speech tier |
| Rate | Credits: ~$0.038 per minute at Pro, ~$0.028 at Scale | $60 per 1M characters (speech-2.8-turbo); $100 per 1M (speech-2.8-hd), international |
| Second price list | None — one entity, one currency | Mainland platform bills in CNY per 10,000 characters — a different unit entirely |
| Commercial use | Prohibited by default; permitted by subscription tier, named only on a pricing card | Not conditioned on tier; ownership and use are addressed in one clause |
| Output ownership | Disclaimed, but §7.1 declines to warrant that you own it | International: “you retain your ownership rights in Client input and generated content”. Mainland: conditional on local law |
| Training on your data | On by default; irrevocable, perpetual, sublicensable. Opt-out is a manual form, forward-only | Permitted by default. No opt-out documented on either platform |
| Voice cloning | Instant from $5/month; professional from $49/month | On the standard API — but cloned voices expire unless used in synthesis |
| Model lifecycle | Alias-level sunsets after October 20, 2026 | No dated end-of-life for any speech model on either platform |
| Speech-to-text | Yes, on the same account | None. No ASR model exists |
| Consent enforcement | Warranty in the policies; no attestation in the cloning API | Identity-gated on the developer; no voice-owner consent requirement published |
| Audio quality | Not compared. Our listening benchmark has not run, so this page assigns no naturalness verdict to either provider. | |
On this page
Winners by category
Every winner below is a documentation-based judgment from published prices, terms and feature documentation. None is a test result, and the categories that require listening have no winner.
| Category | Winner | Why |
|---|---|---|
| Entry price | Cartesia | $5/month with commercial rights and cloning; MiniMax has no documented free speech tier |
| Price at volume | MiniMax | Published per-character rates suit long-form batch work with a dedicated asynchronous endpoint |
| Model stability | MiniMax | No dated end-of-life for any speech model; Cartesia retires two aliases inside seven weeks |
| Data terms | Cartesia, narrowly | Both take broad rights, but at least Cartesia documents an opt-out route. MiniMax documents none |
| Ownership clarity | MiniMax | Its international terms grant ownership outright; Cartesia declines to warrant it |
| Legal simplicity | Cartesia | One entity, one currency, one governing law. MiniMax runs two of each |
| Voice-agent completeness | Cartesia | Speech-to-text on the same account; MiniMax publishes no ASR model at all |
| Cloning lifecycle | Cartesia | Clones persist; MiniMax voices are temporary unless used in synthesis |
| Pricing transparency | MiniMax | Real server-rendered tables per model. Cartesia publishes credits and approximate minutes |
| Developer tooling | Cartesia | MiniMax publishes no branded speech SDK in any language |
| Naturalness | No winner | Requires listening. Our benchmark has not run |
| Pronunciation and difficult text | No winner | Requires listening. Our benchmark has not run |
Price and units
Neither of these can be laid on the other’s axis without care, because they meter differently.
Cartesia sells credits and publishes its own credits-to-minutes approximations: Pro at $5 for roughly 133 minutes works out near $0.038 per minute, Scale at $299 for roughly 10,667 minutes near $0.028. Those are our computations from the vendor’s own figures, and the per-minute rate barely improves across a sixty-fold jump in volume — what the higher plans buy is concurrency and headroom, not a materially better rate.
MiniMax sells characters, at $60 per million on speech-2.8-turbo and $100 per million on speech-2.8-hd, with a dedicated asynchronous endpoint for long text at the same rates. A 100,000-character chapter therefore costs $6.00 on the turbo model. That is arithmetic from list prices, not a measurement.
MiniMax also has a second price list you must not confuse with the first. Its mainland platform prices the same models in yuan per ten thousand characters — a different currency and a different unit. We publish both as the vendor states them and perform no conversion, because converting would produce a number MiniMax does not publish.
One further MiniMax caveat: its international audio subscription is denominated in “audio points”, and the documentation never defines what an audio point is. No subscription tier can therefore be normalized to a per-character cost. Budget on the pay-as-you-go rates.
Data rights and ownership
Both vendors take broad rights over what you send and what you get back. The difference is whether there is a documented way out.
Cartesia’s clause 5.3(c) grants “a non-exclusive, irrevocable, perpetual, worldwide, royalty-free, fully paid, transferable, sublicensable right and license to use any Inputs and Outputs…” The opt-out under 5.3(d) exists but is a manual form, scoped by category, and forward-only — it “will not affect any uses of Your Content (or improvements to the Models as a result) prior to that date.” No dashboard toggle is documented on any plan.
MiniMax’s international terms address ownership and use in a single clause: “As between you and us, and to the extent permitted by applicable laws, you retain your ownership rights in Client input and generated content. We may use the input and generated content to provide, maintain, develop, and improve our Services, comply with applicable law, enforce our terms and policies, and keep our Services safe.” Read both sentences together: the first grants ownership, the second reserves a broad use right that is not qualified by anonymisation or by plan. We found no opt-out documented on either MiniMax platform.
On ownership the positions invert. MiniMax grants it outright on the international platform, and conditions it on local law on the mainland one. Cartesia disclaims ownership of your outputs in clause 5.3, then states in clause 7.1 — capitals in the original — that “THE CARTESIA ENTITIES DO NOT REPRESENT OR WARRANT THAT YOU ARE THE LEGAL OWNER OF ANY OUTPUT…”
What this means for you: if your requirement is that a vendor does not train on your content, neither of these providers satisfies it out of the box, and only one gives you a documented route to ask.
Two kinds of complexity
Cartesia’s risk is on the calendar. sonic-2 and sonic-turbo stop working after October 20, 2026, at alias level — all snapshots — so the usual defence of pinning a dated version does not apply. It has already executed sunsets in June, and its Voice Changer endpoints passed their August 20 sunset with no documented replacement while the marketing page stayed live. Its Terms, meanwhile, are dated June 14, 2024 and contain no model names and no lifecycle language at all, so every one of those commitments is documentation-level rather than contractual.
MiniMax’s risk is jurisdictional. Its international platform contracts through a Singapore entity under Singapore law with binding arbitration — no court option and no small-claims carve-out — and states that personal data is “stored in the data center located in the United States”. Its mainland platform contracts through a Chinese entity under PRC law with litigation. A price comparison between the two properties is not like-for-like, and neither is a legal review.
MiniMax carries one further operational trap: cloned and designed voices are temporary. Its documentation states that “Voices produced by cloning and voice design are temporary: the fee is charged only on first use in speech synthesis”. Create a clone in a batch and synthesize later, and the gap is where you lose voices. Cartesia’s clones persist, and carry across its model migrations with the same voice IDs.
On the other side of the ledger, MiniMax has the better lifecycle story: no dated end-of-life exists for any of its speech models on either platform. We checked for one specifically.
What each can actually build
If you are building a voice agent, this section decides it.
Cartesia sells both halves. Sonic for synthesis and Ink for transcription on the same account, with prepaid agent allowances on every plan. Concurrency is the limit to design around — 2 on Free, 3 on Pro, 5 on Startup, 15 on Scale — and no requests-per-second figure or SLA is published. Note also that agent calls are recorded by default, with no per-call switch documented to turn it off mid-call.
MiniMax sells one half. Its complete international documentation index contains no ASR, transcription or speech-to-text page and no such model. MiniMax can speak but cannot listen, so a voice agent needs a second supplier. It also publishes no branded speech SDK in any language — integration is direct HTTP — and two model identifiers, speech-01-hd and speech-01-turbo, are accepted by every speech endpoint while appearing in no models table and carrying no published price.
One obligation MiniMax places on you that Cartesia does not: an affirmative AI-disclosure duty in its international terms. Set against that, the international synthesis endpoint exposes no watermark parameter at all, though the mainland one does. The duty to label is yours, and the tooling to do it in the audio is not provided on the platform most non-Chinese buyers will use.
Voice quality, pronunciation and naturalness
We have not compared how these two providers sound, and we will not imply otherwise.
Our standardized listening benchmark is designed but not running. Until it does, this page carries no naturalness verdict, no pronunciation comparison and no scores for either provider. Both vendors publish quality claims of their own; those belong to them, and we do not repeat them as findings.
Choose Cartesia if · Choose MiniMax if
Choose Cartesia if:
- You are building a voice agent. Speech-to-text on the same account is decisive; MiniMax has no ASR at all.
- You want the cheapest commercial entry. $5 a month with a commercial licence and instant cloning.
- Your cloned voices must persist. MiniMax’s expire unless used in synthesis.
- One jurisdiction is a requirement. Cartesia is a single US entity; MiniMax is two entities under two legal systems.
Choose MiniMax if:
- Long-form batch synthesis is the job. Published per-character rates and a purpose-built asynchronous endpoint.
- You need a model that will still be there next year. No dated end-of-life exists for any MiniMax speech model.
- You want ownership granted outright rather than merely not claimed.
- You want prices in real tables per model rather than credits and approximations.
Sources and what we could not verify
Every figure on this page was read from the two vendors’ official sources on September 1, 2026. Full detail is in the individual reviews: Cartesia review and MiniMax review. How we verify anything is described on the methodology page.
| Source | Used for |
|---|---|
| cartesia.ai/pricing | Plans, credits, minute approximations, concurrency, cloning gates |
| Cartesia Terms of Service | Clauses 4.1, 5.3, 5.3(c), 5.3(d), 7.1 (Last Revised June 14, 2024) |
| Cartesia API changes | Sunset tables, scope and error behaviour |
| Cartesia agent observability | Default call recording |
| MiniMax international pricing | USD speech rates and units |
| MiniMax mainland pricing | CNY rates and the per-10,000-character unit |
| MiniMax API overview | Cloning lifecycle, charging on first use, identity gating |
| MiniMax international Terms | Entity, ownership and use clause, disclosure duty, arbitration (Effective March 30, 2026) |
| MiniMax documentation index | Full-index checks for ASR coverage and the audio-point definition |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- How either provider sounds. No listening test has been run.
- What a MiniMax “audio point” is. Used to price the audio subscription and defined nowhere.
- Whether any MiniMax opt-out from model-development use exists. We found none documented on either platform; we record its absence from the pages read.
- What
speech-01-hdandspeech-01-turbocost or do. Accepted by every MiniMax speech endpoint, documented nowhere. - The exact Cartesia credits-to-characters conversion. It publishes credits and approximate minutes only.
- Whether Cartesia’s Voice Changer endpoints are actually off. Their sunset date passed while the capability page stayed live.
- Any latency figure for either provider in a form we would publish.
Change log
September 2, 2026 — First publication as a documentation-based comparison. All figures read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · When our audio benchmark runs, this page gains a head-to-head listening test using the same prompts and documented conditions for both providers, and the naturalness and pronunciation rows gain real winners (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.