MiniMax Review 2026

✓ Docs-verified · Sep 1, 2026◌ Not audio-tested△ What we could not verify

Published September 1, 2026 · Every fact below was read from MiniMax’s own developer documentation, pricing tables and platform agreements on September 1, 2026, with the source linked beside the claim. Where a document required an unusual reading method, this review says so.

Quick verdict

MiniMax sells capable speech synthesis through two platforms that are not the same product in any way that matters to a contract. The international platform bills in US dollars at $60 per million characters for speech-2.8-turbo; the mainland platform bills in yuan per ten thousand characters. They run under different corporate entities, different governing law and different dispute mechanisms (international pricing, mainland pricing, both checked September 1, 2026).

The operational trap is a seven-day timer. Cloned and designed voices are temporary by default: “Voices produced by cloning and voice design are temporary: the fee is charged only on first use in speech synthesis” — and a voice that is never used in synthesis expires (API overview, checked September 1, 2026). A clone you create and do not immediately use is a clone you may not have.

And you carry a disclosure duty. MiniMax’s international terms impose an affirmative obligation on the customer around labelling AI-generated content — while the international synthesis endpoint exposes no watermark parameter at all, though the mainland one does. See licensing.

Verdict basis: official documentation only, checked September 1, 2026 — no account, no API call, no listening test, no numeric score (how we verify).

MiniMax at a glance — verified September 1, 2026
Best forDevelopers who want low-cost long-form synthesis and can accept Singapore-law arbitration and a default training licence
Not best forAnyone needing speech-to-text from the same vendor, a native SDK, an opt-out from model training, or a defined subscription billing unit
Starting pricePay-as-you-go, no free speech tier documented. International: $60 per 1M characters (speech-2.8-turbo), $100 per 1M (speech-2.8-hd). Mainland prices in CNY per 10,000 characters
Voice cloningAvailable, identity-gated on the developer. No voice-owner consent requirement is published in the developer documentation
Output ownershipInternational: “you retain your ownership rights in Client input and generated content”. Mainland: conditional on what local law provides
Training on your dataPermitted by default, with no opt-out documented on either property
Test statusDocumentation-verified. Our listening benchmark has not run, so there is no independent voice-quality score and no measured latency on this page.

Every price, licensing term and feature claim below was read from MiniMax’s official pages on September 1, 2026, with the source linked beside the fact. Our standardized audio benchmark is designed but not yet running, so this review contains no listening-test claims, no naturalness verdicts and no scores.

Strongest advantages:

Important disadvantages:

On this page

Two platforms, two legal regimes

This comes first because almost every other fact on the page depends on which platform you are actually buying from.

The two MiniMax platforms — verified September 1, 2026
InternationalMainland
Platformplatform.minimax.ioplatform.minimaxi.com
Contracting entityNanonoble Pte Ltd.上海稀宇科技有限公司
Governing lawSingaporePRC
DisputesBinding SIAC arbitration — no court option, no small-claims carve-outLitigation
Currency“all payments are in USD”CNY
Billing unitPer 1 million charactersPer 10,000 characters (单价 元/万字符)
Personal data location“stored in the data center located in the United States”Not stated in the same terms
Output ownershipGranted flatlyConditional on what local law provides

What this means for you: a price comparison between the two properties is not a like-for-like comparison, and neither is a legal review. If your procurement team approved the international terms, the mainland platform is a different contract with a different forum and a different remedy path. Choose deliberately, and record which one you are on.

Who should use it — and who should look elsewhere

Choose MiniMax if:

Look elsewhere if:

Current price, free plan, commercial rights

The short version: pay-as-you-go pricing is published and cheap, the subscription unit is undefined, and there is no documented free speech tier.

Published speech prices, exactly as displayed — verified September 1, 2026
PlatformModelPrice as displayedUnit as the vendor words it
Internationalspeech-2.8-turbo$60per 1M characters
Internationalspeech-2.8-hd$100per 1M characters
Mainlandspeech-2.8-turbo2元/万字符 (CNY per 10,000 characters)
Mainlandspeech-2.8-hd3.5元/万字符 (CNY per 10,000 characters)

Read from the “Audio” section of the international pay-as-you-go page and the 语音 section of the mainland equivalent, both checked September 1, 2026. Both are server-rendered tables rather than JavaScript-injected values. We publish both units as the vendor states them and perform no currency conversion; converting would produce a number MiniMax does not publish.

The same asynchronous long-text endpoint is priced at the same rates on both properties for the respective models.

The subscription unit is never defined

The international platform sells an audio subscription whose tiers are denominated in “Audio points per month” — 100,000, 300,000, 1,100,000, 3,300,000 and 20,000,000. What an audio point is is not stated. The phrase occurs exactly once across the complete international documentation index we read, and never in a definition.

What this means for you: you cannot compute the effective per-character cost of any subscription tier from published information. Until MiniMax defines the unit, the pay-as-you-go rates are the only figures a buyer can actually budget from — and we would build a projection on those, not on points.

One documented exclusion is worth knowing before you buy a plan for cloning work: the token plan’s own footnote states that coverage spans the model lineup including speech, but that a small number of special models are excluded — and voice design and rapid voice cloning fall outside it. A subscription is not a route to cheap cloning here.

Commercial rights and ownership

The international terms, effective March 30, 2026, address ownership and training in a single clause of the intellectual-property section. Verbatim:

“As between you and us, and to the extent permitted by applicable laws, you retain your ownership rights in Client input and generated content. We may use the input and generated content to provide, maintain, develop, and improve our Services, comply with applicable law, enforce our terms and policies, and keep our Services safe.”

Read the two sentences together, because they are doing different jobs. The first grants you ownership. The second reserves a broad use right over the same material — including to “develop” and “improve” the services — and it is not qualified by anonymisation, by plan, or by anything else. The mainland document words ownership conditionally instead, making it contingent on what local law provides.

We found no opt-out from that use on either property. The phrase does not appear in the documents we read.

What this review covers

This review covers MiniMax’s speech synthesis offering on both the international and mainland platforms: the models and their identifiers, the published pay-as-you-go prices in their own units, the subscription structure, cloning and voice design including their lifecycle, the rate limits, and the terms governing ownership, training, disclosure and disputes.

It does not evaluate how any MiniMax voice sounds. We created no account, made no API call and generated no audio. Rows a review of ours normally carries once our benchmark runs — Tested, Model tested, Voice tested, Benchmark cost — are absent rather than filled with invented values.

A note on how the legal documents were read, because it affects how much weight to give them. Unlike the pricing tables, which are ordinary server-rendered HTML, MiniMax’s platform agreement pages are client-side rendered applications: the served HTML contains almost no visible text, and the full agreement sits inside the page’s data payload. We read the agreement text out of that payload rather than from a rendered browser view. It is unambiguously the vendor’s own served document, and a person opening the URL in a browser sees the same text — but we would rather tell you how we obtained it than imply we read it the ordinary way.

Benchmark audio status

Samples are pending. Our standardized audio benchmark is designed but not running, so this page publishes no audio and no empty player (methodology). What we verified instead is documentation: prices in both currencies and units, model enums on every speech endpoint, cloning lifecycle and gating, rate limits on both properties, and the contract clauses quoted below, all read on September 1, 2026.

Assessment by dimension

No numeric scores appear below. The judgments are qualitative and drawn from what MiniMax documents.

Documentation-based assessment — all sources checked September 1, 2026
DimensionWhat the documentation supports
Pay-as-you-go priceStrong. Published openly in real tables on both properties, and low for long-form work.
Subscription clarityWeak. The billing unit is undefined, so tiers cannot be normalized.
Legal clarityWeak. Two entities, two laws, two dispute mechanisms, and ownership worded differently by property.
Data privacyWeak by default. Broad use right over inputs and outputs with no documented opt-out.
Cloning accessStrong on availability, weak on lifecycle. On the standard API, but voices expire unless used.
Consent controlsAbsent from the product. Identity gating applies to the developer; no voice-owner consent requirement is published in the developer docs.
Developer surfaceWeak. No native speech SDK; no speech-to-text; two undocumented model identifiers accepted by the enums.
Voice qualityNot assessed. We have run no listening test.
LatencyNot assessed and not published. No figure exists for any speech model on either property.

Models, including two that are not documented

The current speech line is speech-2.8-hd and speech-2.8-turbo, priced on both properties. Earlier generations — speech-2.6-hd, speech-2.6-turbo, speech-02-hd and speech-02-turbo — sit inside a collapsed documentation block titled “Legacy Models” on the international property, and its Chinese equivalent on the mainland one, while remaining priced on the current pay-as-you-go page.

No dated end-of-life exists for any speech model on either property. That is a genuine and useful contrast with several competitors in this comparison, and we checked for it specifically. Note that an adjacent audio product — the music and lyrics APIs — does carry a dated service-adjustment notice, but it does not apply to speech.

Two identifiers deserve a warning. speech-01-hd and speech-01-turbo are accepted values in the model enum of every speech endpoint on both properties — synchronous, WebSocket, bidirectional, asynchronous and voice cloning — yet they appear in no models table and carry no published price. An API that accepts a model it does not document is an API that can bill you for something you cannot look up. Do not send either identifier without confirming what it costs.

Voice quality and naturalness

We will not tell you how MiniMax sounds. No listening test has been run for this review.

Nor is there a vendor figure to quote: we found no measured quality or latency number for any MiniMax speech model on either property, across every speech guide, both API references, the API overview, both release-notes pages, the audio FAQ and both pricing pages. That is unusual — most vendors publish at least a marketing benchmark — and it means this section has nothing in it until our own benchmark runs.

Voice cloning, consent, and the seven-day timer

Cloning and voice design are both available on the standard API, which is a real advantage over providers that gate cloning behind enterprise plans. Two documented behaviours change how you should use them.

Voices are temporary until used. The API overview states verbatim: “Voices produced by cloning and voice design are temporary: the fee is charged only on first use in speech synthesis (previews within those APIs do not count)”. The practical effect is a timer: create a clone, fail to use it in synthesis, and it goes away. If your workflow clones in a batch and synthesizes later, that gap is where you lose voices.

Consent is not enforced by the product. Cloning is identity-gated on the developer — the account holder must be verified — but we found no voice-owner consent requirement published anywhere in the developer documentation. The cloning flow takes an uploaded audio file and returns an identifier; it does not ask you to attest that the speaker agreed.

That places MiniMax at the permissive end of this market. Speechify’s API, by contrast, requires the speaker to record a phrase the vendor issues, with no unattended path (consent docs, checked September 1, 2026). If your compliance position depends on the platform enforcing consent, it does not here — and note that you also carry a disclosure duty, described in licensing.

API, SDKs, rate limits

There is no MiniMax speech SDK. The platform’s only SDK page is titled “Integrate via SDK” and its subtitle directs developers to use a third party’s SDK to call a MiniMax language model — not a speech one. Speech integration is direct HTTP.

There is no speech-to-text model. The complete international documentation index contains no ASR, transcription or speech-to-text page and no such model. If you are building a voice agent, MiniMax can speak but cannot listen, and you will need a second vendor for the other half.

Published rate limits for speech on the international property are 60 requests per minute for text-to-audio, 60 for voice cloning and 20 for voice design. The mainland ceiling is documented lower. We found the rate-limit documentation internally inconsistent in places and do not attempt to reconcile it here; treat the published figures as the floor of what to design for and confirm against your own account.

Generation speed and latency

No latency figure of any kind is published for any MiniMax speech model on either property, and we have measured none. This section has nothing to report — not even a vendor claim to quote — and will carry measured figures when our benchmark runs.

Pricing and normalized cost

What you pay, and what you get

On the international platform the rate is the effective cost, because billing is per character with no subscription required: cost = (characters ÷ 1,000,000) × the model rate. A 100,000-character audiobook chapter comes to $6.00 on speech-2.8-turbo and $10.00 on speech-2.8-hd, at the prices published on September 1, 2026. That is arithmetic from list prices, not a measurement of a workload we ran.

The important limitation: the subscription cannot be normalized at all

Because “audio points” is undefined, there is no way to compute what a subscription tier costs per character, and therefore no way to know whether it beats pay-as-you-go. That is not a gap in our research — we looked across the complete documentation index — it is a gap in the published information. Budget on the pay-as-you-go rates.

Who gets the best value

Commercial usage, licensing, privacy

Ownership, and the use right beside it

The international clause is quoted in full in commercial rights above. Its structure is worth restating plainly: you retain ownership of input and generated content, and MiniMax reserves a broad right to use that same content to develop and improve its services. Both halves are in one clause and there is no separate output-licence clause anywhere in the agreement.

The mainland document words ownership conditionally — contingent on what applicable law provides — rather than granting it outright. Same company, two different formulations, and which one applies depends on which platform you signed up to.

The same intellectual-property section places warranties on you regarding the inputs you supply, with an indemnity tail. If you are cloning a voice, that is the clause that allocates the risk of having got consent, and it allocates it to you.

You carry an AI-disclosure duty

MiniMax’s international terms impose an affirmative obligation on the customer concerning the labelling of AI-generated content. This is not a courtesy request buried in a policy; it sits in the user-obligations section of the agreement effective March 30, 2026, and it applies on the Singapore-entity international terms, not only on the mainland ones.

Set against that, the international synthesis endpoint exposes no watermark parameter at all. We enumerated every property of the international text-to-audio request object — model, text, stream, stream options, voice settings, audio settings, pronunciation and the rest — and none of them controls a watermark. The mainland endpoint does expose one.

So the obligation to label is yours, and on the international platform the tooling to do it in the audio itself is not provided by the API. Plan your disclosure mechanism outside the request.

Disputes and data location

The international terms select Singapore law with binding SIAC arbitration — no court option and no small-claims carve-out — and state that personal data is “stored in the data center located in the United States”. The mainland terms select PRC law and litigation. For most buyers outside China the international regime is the relevant one, and the arbitration clause is the sentence to take to counsel before signing.

Pros

Cons

Best use cases

Where MiniMax’s documented strengths land — editorial judgment from documentation, not test results (September 1, 2026)
Use caseFitWhy
Long-form batch narrationStrongLow per-character price and a purpose-built asynchronous endpoint
Cost-sensitive high volumeStrong on paper$60 per million characters on the turbo model, published openly
Cloning-based productsMixedAvailable on the standard API, but excluded from the token plan and subject to expiry
Real-time voice agentsWeakNo speech-to-text model exists; you need a second vendor for the listening half
Regulated deploymentWeakDefault training use with no opt-out; arbitration-only disputes; ownership worded differently by property
Teams needing a native SDKWeakNone exists for speech in any language

And the inverse, because a recommendation without one is not worth much: if your organisation needs to point at a single contract, in a single currency, with a defined billing unit and an opt-out from model training, MiniMax does not offer that on either of its platforms today.

Alternatives

Organized by the reason you would leave.

Direct comparison links

Head-to-head pages pairing MiniMax against individual competitors are on this site’s roadmap but not published yet, and we do not link to pages that do not exist. Until they are live, the flagship ranking carries side-by-side pricing, licensing and feature tables for all ten providers we cover.

Final verdict

Choose MiniMax when the job is long-form synthesis, the budget is tight, and you are comfortable with the international terms. The pay-as-you-go rates are genuinely low, the asynchronous long-text endpoint is the right tool for audiobook-shaped work, cloning is available without a sales call, and — unusually in this comparison — no speech model carries a published end-of-life date hanging over it.

Do not choose it when your project needs the listening half of a voice agent, an opt-out from model training, a subscription you can price, or a single legal regime. The two-platform split is not cosmetic: different entities, different law, different dispute mechanism, different currency, different billing unit, and ownership granted flatly in one document and conditionally in the other. Add a cloning lifecycle that quietly expires unused voices and a disclosure duty with no watermark parameter to satisfy it, and the low headline price is buying you a set of obligations to manage.

What would change this verdict is sound — and, before that, a definition. Publishing what an audio point is would make the subscription comparable overnight; documenting the two orphaned model identifiers would close a billing risk. When our audio benchmark runs, this page gains measured comparisons and the assessment table gains scores that mean something. Until then this is a verdict about two rate cards and two contracts.

Frequently asked questions

Did you actually test MiniMax’s voices?

No. We created no account, made no API call and generated no audio for this review. Every statement here comes from MiniMax’s own documentation, read on September 1, 2026. Our listening benchmark is designed but not running (methodology).

How much does MiniMax text-to-speech cost?

On the international platform, $60 per million characters for speech-2.8-turbo and $100 per million for speech-2.8-hd. The mainland platform prices the same models in yuan per ten thousand characters. We publish both units as stated and perform no conversion.

What is an “audio point”?

MiniMax does not say. The subscription tiers are denominated in audio points per month, but the term is never defined in the documentation we read, so no subscription tier can be normalized to a per-character cost.

Do my cloned voices last?

Not automatically. The documentation states that voices produced by cloning and voice design are temporary and that the fee is charged only on first use in synthesis. A clone you never use in synthesis expires.

Does MiniMax require consent to clone a voice?

Cloning is identity-gated on the developer, but we found no voice-owner consent requirement published anywhere in the developer documentation. The terms place warranties about your inputs on you, with an indemnity obligation attached.

Does MiniMax train on my data?

The international terms permit MiniMax to use input and generated content to provide, maintain, develop and improve its services. That use is not qualified by plan or by anonymisation, and we found no opt-out documented on either platform.

Can MiniMax transcribe audio as well as generate it?

No. There is no ASR or speech-to-text model in the documentation. A voice agent built on MiniMax needs a second vendor for transcription.

Which platform should I use?

They are different contracts, not different endpoints. The international platform runs under a Singapore entity with Singapore law and binding arbitration; the mainland platform under a Chinese entity with PRC law and litigation. Decide deliberately and record which one you are on.

Sources and what we could not verify

Every changing fact on this page was read from official MiniMax sources on September 1, 2026. Prices come from server-rendered tables. Agreement text was read from the served page payload, as explained in what this review covers. Where the two properties differ, both are published and neither is treated as the default.

Official sources consulted — all checked September 1, 2026
SourceUsed for
International pay-as-you-go pricingUSD speech rates, units, asynchronous endpoint pricing
Mainland pay-as-you-go pricingCNY speech rates and the per-10,000-character unit
Token planCoverage footnote and the exclusion of voice design and rapid cloning
API overviewCloning and voice-design lifecycle, charging on first use, identity gating
Text-to-audio HTTP referenceRequest object properties — used to establish the absence of a watermark parameter
Models introductionCurrent and legacy speech model lineup
Rate limitsPublished requests-per-minute figures for speech endpoints
Integrate via SDKEstablishing that no MiniMax speech SDK is published
Documentation indexFull-index checks for ASR coverage and the “audio point” definition
International Terms of ServiceEntity, ownership and use clause, user obligations, disclosure duty, arbitration, data location (Effective March 30, 2026)

What we could not verify

Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026, recorded rather than guessed:

Change log

September 1, 2026 — First publication. All prices, terms, model identifiers, limits and feature statements read from official MiniMax sources on this date, across both the international and mainland platforms.

Published September 1, 2026 · Benchmark audio, measured costs and scores will be added with a dated entry here when the audio benchmark runs (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.