AI Voice Generator Pricing Comparison
Published September 2, 2026 · Every rate on this page was read from the vendor’s own pricing page, documentation or public price API on September 1, 2026. This is a price comparison built from published rates, not an audio test — nothing here says anything about how any of these voices sound.
The short answer
One million characters of synthesis costs $4 at the floor. Published metered rates run from there to $160; subscription plans, converted at full allowance use, reach $200. That is a fortyfold spread on published rates and a fiftyfold spread once subscriptions are included — and it is the single most important fact about this market’s pricing. The two tables below keep those apart, because one is what vendors publish and the other is our arithmetic.
You cannot see that spread from the vendors’ own pages, because they bill in four incompatible units: per character, per credit, per token, and per hour of generated audio.
Three of the ten do not render a readable text-to-speech price on the page a buyer would visit — Murf, Azure and Resemble AI. Two more publish a figure that is real but incomplete or hard to reach. Murf’s two pricing pages serve an empty container with no dollar sign anywhere in the HTML. Azure’s renders a dash where the price belongs, 142 times. Resemble AI’s never mentions speech synthesis at all.
What this page will not tell you is whether the expensive end is worth it. That is an audio question, and our listening benchmark is designed but not running. A fiftyfold price difference is a documented fact; a fiftyfold quality difference is not.
Basis: published vendor rates only, checked September 1, 2026 — no accounts created, no invoices generated (how we verify).
On this page
The unit problem
Comparing these ten providers on price is not a matter of lining up numbers. They do not measure the same thing.
- Per character. Google Cloud, Amazon Polly, Azure, MiniMax, Speechify’s API and Murf’s API all bill by characters of input text. These are directly comparable.
- Per credit. ElevenLabs and Cartesia sell a monthly credit allowance. Both publish a credit-to-character factor, so both can be converted — but only because they chose to publish it.
- Per token. OpenAI’s newer speech model is billed in audio tokens. No conversion to characters or minutes is published anywhere on its pricing page.
- Per hour of generated audio. Murf’s Studio plans are sold in hours of voice generation. Text length has nothing to do with it, and no characters-per-hour factor is published.
The fourth unit is the one that breaks the comparison outright. An hour of audio corresponds to some number of characters only if you assume a speaking rate, and assuming a speaking rate to produce a price is inventing a number. We do not do that. Murf’s Studio plans appear on this page in their own unit and are not ranked against anything.
Every rate, in one unit
Below is every published text-to-speech rate we could verify, expressed as US dollars per one million characters. Rows marked as published needed no arithmetic: the vendor states the rate in this unit or a directly equivalent one. Rows marked vendor factor were converted using a conversion the vendor itself publishes, quoted in the notes.
| Provider | Model or tier | $ / 1M characters | Basis |
|---|---|---|---|
| Google Cloud | Standard voices | $4.00 | As published — “US$0.000004 per character” |
| Amazon Polly | Standard voices | $4.00 | As published, commercial regions |
| Speechify | API, Scale overage rate | $6.00 | As published, on a $499/month plan |
| Speechify | API, Pro overage rate | $8.00 | As published, on a $99/month plan |
| Murf AI | API, Falcon | $10.00 | As published — “$0.01/1000 characters” · see the unit note below |
| Speechify | API, Starter overage rate | $10.00 | As published, on a $10/month plan |
| Azure Speech | S1 Neural | $15.00 | As published, in the public price API — not on the pricing page |
| OpenAI | tts-1 | $15.00 | As published — “$15.00 / 1M characters” |
| Google Cloud | Neural2 and Polyglot | $16.00 | As published |
| Amazon Polly | Neural voices | $16.00 | As published |
| Azure Speech | Neural HD | $22.00 | As published, in the public price API |
| Amazon Polly | Generative voices | $30.00 | As published |
| OpenAI | tts-1-hd | $30.00 | As published — “$30.00 / 1M characters” |
| Murf AI | API, Gen 2 | $30.00 | As published — “$0.03/1000 characters” · see the unit note below |
| Cartesia | Scale overage rate | $38.00 | Vendor factor — “$38 per 1M credits”, with TTS at “approximately 1 credit per character” |
| Cartesia | Startup overage rate | $45.00 | Vendor factor — “$45 per 1M credits” |
| Google Cloud | Instant custom voice | $60.00 | As published |
| MiniMax | speech-2.8-turbo | $60.00 | As published — “$60/M characters”, international platform |
| Cartesia | Pro overage rate | $65.00 | Vendor factor — “$65 per 1M credits” |
| Amazon Polly | Long-Form voices | $100.00 | As published |
| MiniMax | speech-2.8-hd | $100.00 | As published — “$100/M characters”, international platform |
| Google Cloud | Studio voices | $160.00 | As published |
The unit note on Murf. Murf’s help centre states “$0.01/1000 characters” for Falcon and “$0.03/1000 characters” for Gen 2, and a third Murf page independently states “Falcon 2 costs 1 cent per 1000 characters.” But Murf’s own pricing payload sets charactersPerUnit to 10,000 for pay-as-you-go, while the page code labels the same rates “per 1,000 characters”. Read against 10,000 characters the payload reconciles with the help centre; read against the label the page code attaches, the same numbers would render ten times higher. We publish the help-centre figure, which two Murf sources agree on, and we do not resolve the discrepancy by arithmetic. It is Murf’s to resolve.
The Azure note. Azure’s $15 and $22 figures do not come from its pricing page, which shows no numbers at all. They come from the Azure Retail Prices API, which returns retailPrice 15.0 per 1M for the S1 Neural meter and 22.0 per 1M for Neural HD, the latter with an effective date of March 1, 2026. Microsoft’s own quotas documentation independently states “Multiply the result by the unit price of $15 per million characters” — though that sentence sits inside a worked example for the Standard tier rather than on a rate card, and Azure’s pricing page carries the caveat “Prices are estimates only and are not intended as actual price quotes.”
The subscription rates, and why they are best cases
ElevenLabs and Cartesia sell monthly credit allowances rather than metering usage. Dividing the plan price by the included credits gives an effective rate — but only if you consume the entire allowance. Use half of it and your real cost per character doubles. Every figure in this table is therefore a best case, and it is our arithmetic from two vendor figures, not a rate the vendor publishes. Both inputs are shown so you can check the division yourself.
| Provider | Plan | Monthly price | Included credits | Effective $ / 1M characters |
|---|---|---|---|---|
| Cartesia | Scale | $299 | 8,000,000 | $37.38 |
| Cartesia | Startup | $49 | 1,250,000 | $39.20 |
| Cartesia | Pro | $5 | 100,000 | $50.00 |
| ElevenLabs | Pro | $99 | 600,000 | $165.00 |
| ElevenLabs | Business | $990 | 6,000,000 | $165.00 |
| ElevenLabs | Scale | $299 | 1,800,000 | $166.11 |
| ElevenLabs | Creator | $22 | 121,000 | $181.82 |
| ElevenLabs | Starter | $6 | 30,000 | $200.00 |
The credit-to-character factor is the vendor’s own in both cases. ElevenLabs: “For V2 Multilingual models, 1 text character equals 1 credit.” Cartesia: “Standard TTS costs approximately 1 credit per character. The exact number of credits can vary slightly due to transcript pre-processing.” Read that second sentence before relying on the third decimal place of anything here.
One material qualifier on the ElevenLabs column. ElevenLabs documents a second rate for its faster model: “0.5 and 1 credit per character” depending on model, with Flash at the lower end. Generate on Flash rather than V2 Multilingual and every ElevenLabs figure above halves — $200 becomes $100, $165 becomes $82.50. Which model you use changes the effective price by a factor of two, and the plan card does not tell you that.
Cartesia’s published overage rates in the previous table are cleaner evidence than the subscription arithmetic, because they are true marginal rates the vendor states per million credits. When they disagree slightly with the plan division — $38 against $37.38 on Scale — that gap is real, not a rounding artefact: the allowance is priced marginally better than the overage.
What we refused to convert
Four things on this market’s price lists cannot be expressed in characters without inventing a number. We name them rather than estimating them.
Murf’s Studio plans are sold in hours. Creator Lite is $29/month for 2 hours of voice generation; Creator Plus $49 for 4 hours; Business Lite $99 for 8 hours; Business Plus $199 for 20 hours. Converting hours to characters requires a speaking rate, and Murf publishes none. These plans are not comparable to any per-character rate on this page and we do not pretend otherwise.
OpenAI’s gpt-4o-mini-tts is billed in audio tokens — $12.00 per 1M audio output tokens, plus $0.60 per 1M text input tokens. No token-to-character factor and no per-minute estimate is displayed for it anywhere on OpenAI’s pricing page. Its two older siblings, tts-1 and tts-1-hd, carry a per-cell override reading “/ 1M characters”, which is why those two appear in the main table and this one does not.
Resemble AI publishes no speech-synthesis rate on its pricing page at all. We searched the full fetched page: “text-to-speech”, “voice cloning”, “speech synthesis”, “voice generation” and “per character” each return zero occurrences. What the page prices is deepfake detection, billed per second of processed audio, under the title “Deepfake Detection & AI Security Pricing”. The only synthesis prices we could find live in a changelog entry and a product FAQ: voice cloning at $2 per clone with the first clone free, and the Voice Cloning API gated behind “a Business plan or higher”. Since this page was written we have found the answer to the obvious question. (correction record) Resemble AI’s own machine-readable description of itself, dated August 13, 2026 and read on September 2, 2026, states verbatim: “Resemble AI does not sell voice generation products and is not accepting new voice customers. Voice continues as open research only.” The pricing page prices no speech synthesis because none is sold. The changelog and FAQ figures above are retained as a record of what was published; they are not live rates. See the full review.
Cartesia’s voice agents are billed in a second currency. Its own documentation opens: “Cartesia meters model usage in credits and agent usage in agent dollars.” Each plan carries both — the Scale plan gives 8M credits and, separately, $299 to spend on managed agents. Agent minutes are billed per minute in USD at $0.06 for a base call. That is a different product on a different meter, and it does not belong in a per-character table.
The exception that proves the rule. Google’s Gemini-TTS models are billed in audio tokens like OpenAI’s — $10.00 per 1M audio output tokens for Gemini 2.5 Flash TTS, $20.00 for Gemini 3.1 Flash TTS (Preview) and Gemini 2.5 Pro TTS — but Google publishes the missing factor in a footnote: “Audio tokens correspond to 25 tokens per second of audio.” That single sentence makes the rate convertible to time. One million audio tokens is 40,000 seconds, or 11.1 hours, which puts Gemini 2.5 Flash TTS at roughly $0.90 per hour of generated audio and the two $20 models at roughly $1.80 per hour — our arithmetic, from Google’s rate and Google’s own conversion. Note what we still do not do: we do not convert those back into characters, because that would need a speaking rate and nobody publishes one. Every Gemini-TTS row also reads “Not available” under free usage. The difference between this and the OpenAI case is one footnote.
The vendors whose prices you cannot read
This is the finding we did not expect. Five of the ten do not put a readable price in front of a buyer who visits their pricing page.
- Murf AI. Both
murf.ai/pricingandmurf.ai/api/pricingreturn HTTP 200 and serve an empty container — a<div id="root"></div>and a noscript notice. The served HTML of the first contains zero dollar-sign characters. Every Murf figure on this page comes from Murf’s own JSON pricing endpoints, whose paths appear in Murf’s own JavaScript bundle, or from server-rendered help-centre prose. - Azure Speech. The text-to-speech price cells on Microsoft’s speech pricing page render as the literal string
$-. That string occurs 142 times. The only dollar-sign-plus-digits string in the entire 646KB document is “$200”, the free-account credit. - Resemble AI. The pricing page prices a different product, as described above.
- Speechify, consumer app. The annual price is never painted. It exists only inside the page’s embedded framework payload, as
"priceYearly":"$139". The monthly $29 renders; the annual figure is data we read out of the page rather than a price we saw displayed, and we label it that way. - Cartesia. The answers in its pricing-page FAQ — which is where the credit-to-character conversion lives — are not in the rendered document.
None of this is hidden pricing in the sense of a sales gate; the numbers exist and are public. But four of these five require reading a JSON payload, a price API or a JavaScript bundle to see a figure, which is not a reasonable thing to ask of someone comparing options. Where a vendor’s own page and its own payload disagree, we publish both readings rather than choosing one.
The largest multiples inside a single vendor — Google’s 40× from Standard to Studio, Amazon’s 25× from Standard to Long-form — are what vendors charge for a premium tier they describe as more human-sounding. What each of them claims for that premium, in its own words, is set out in most realistic AI voices. No vendor publishes a measurement supporting any of these multiples.
What counts as a character
Two vendors can charge the same headline rate per character and still bill you differently, because they count characters differently. Only some of them say so.
- Google Cloud counts everything you send. Verbatim: “The total number of characters in the input string are counted for billing purposes, including spaces and newline characters. All Speech Synthesis Markup Language (SSML) tags (except the
<mark>tag) are also included in the character count.” Heavily marked-up SSML costs more than the words it wraps. Google also states the friendlier half: for WaveNet and Standard voices a multi-byte character such as Japanese is charged as one character, not as its bytes. - MiniMax counts a Chinese character as two — but publishes the rule only on its mainland platform, in Chinese, verbatim: “1个汉字算2个字符”. The international English pricing page carries no character-counting rule at all. If you are synthesizing Chinese, the effective rate is not the one on the card.
- Amazon Polly charges for metadata requests too. Its Standard, Neural and Long-Form rates read “for speech or Speech Marks requests” — requesting timing metadata bills at the same per-character rate as generating audio. The Generative tier is the only one whose wording omits Speech Marks.
- Cartesia bills structural markup. Verbatim: “Each break tag is counted as 1 credit in the transcript of a TTS request.” Its infill feature costs “a fixed 300 credits, plus the standard TTS rate applied to the infill transcript”. Cartesia also states that failed requests are free: “Credits are only used by successful requests; errors will not consume credits.”
- Azure bills different features on different meters. Verbatim: “Text to Speech: speech synthesis usage is billed per character. Avatar is billed per second. Training and model hosting is billed per second.” A custom voice therefore accrues hosting charges whether or not you synthesize anything with it.
Amazon also publishes a second, higher price list for AWS GovCloud — Standard at $4.80 and Neural at $19.20 per million characters — covering only two engines. If you are procuring for US government workloads the commercial rates above are not your rates.
Free allowances, in the same unit
Free allowances are quoted in the same four incompatible units, so the same normalization applies. These are the published monthly allowances, with the durations exactly as each vendor states them.
| Provider | Free allowance | Duration as stated |
|---|---|---|
| Google Cloud | 4M characters/month, Standard voices | No stated end |
| Google Cloud | 1M characters/month, Neural2 and Polyglot; 1M for Studio | No stated end |
| Amazon Polly | 5M characters/month, Standard voices | No duration stated — the only Polly tier without one |
| Amazon Polly | 1M characters/month, Neural | First 12 months |
| Amazon Polly | 500K Long-Form; 100K Generative | First 12 months |
| Murf AI | 100,000 characters, API free trial | Trial allotment, not monthly |
| Speechify | 50,000 characters/month, API — hard cap, no overage | Ongoing free plan |
| Cartesia | 20,000 credits/month — about 27 minutes by Cartesia’s own estimate | Recurring $0/month plan |
| ElevenLabs | 10,000 credits/month — about 10,000 characters on V2 Multilingual | Ongoing free plan |
| Azure Speech | 500,000 characters/month, Neural, on the F0 tier — plus a $200 new-account credit | Recurring monthly |
| OpenAI | No free text-to-speech allowance found | — |
The two cloud platforms are two orders of magnitude more generous than the creator tools, and the gap is structural rather than promotional: 4 to 5 million characters a month against 10,000 to 50,000. What the cloud allowances require in return is billing setup, and what several creator free tiers withhold is the right to publish what you make. The free tier comparison works through both, and the licensing position is set out in our guide to can you use AI voices commercially.
When the cheapest rate is not the cheapest bill
Four things reliably move a real invoice away from the rate card.
Unused allowance is spent money. Every figure in the subscription table assumes you consume the whole monthly credit balance. Cartesia is explicit that unused credits carry over — “You will keep the credits in your account until you use them” — which softens this considerably; where a vendor does not say that, assume the allowance expires.
The model you pick can double the rate. ElevenLabs charges between 0.5 and 1 credit per character depending on model. Amazon’s Long-Form tier costs 25 times its Standard tier. Google’s Studio voices cost 40 times its Standard voices. These are the same vendor, the same account, the same API call with one parameter changed.
Overage terms differ in kind, not degree. Speechify’s free API plan is a hard cap: “Hard cap — no overages, no surprise bills.” Cartesia lets a paid balance go negative if you enable overages and blocks requests if you do not. A hard cap and an uncapped meter are different risk profiles, not different prices.
Seats and subscription fees sit on top. Murf’s Studio plans bundle a fixed number of users; Resemble’s Team plan is $350/month before any usage; ElevenLabs’ Scale and Business plans include 3 and 10 seats. A per-character rate compared without its subscription fee is not a comparison.
The one honest summary: at the volumes most individual buyers actually generate, the subscription fee dominates and the per-character rate barely matters. At API volume the reverse is true, and the fiftyfold spread at the top of this page becomes the whole story.
Frequently asked questions
Which AI voice generator is cheapest?
By published rate, Google Cloud Standard voices and Amazon Polly Standard voices are tied at $4.00 per million characters, and both offer the largest free allowances in this comparison. Both require billing setup on a cloud account, which is a real barrier if you are not already a cloud customer.
Is a more expensive voice a better voice?
This page cannot answer that, and neither can any page on this site yet. Our listening benchmark is designed but not running, so we have run no test that would let anyone rank these providers on how they sound. The fiftyfold price spread is a documented fact about rate cards, nothing more.
Why can I not just compare the prices on the vendors’ own pages?
Because they are quoted in four different units, and five of the ten do not render a readable price on the page you would visit. Murf serves an empty container, Azure renders a dash, and Resemble’s pricing page prices a different product entirely.
How much does one hour of audio cost?
We do not publish that, because converting characters to minutes requires assuming a speaking rate and no vendor publishes one for billing. Cartesia is the exception — it states “One minute of audio generation requires 750-800 credits” — and that figure is Cartesia’s alone; it does not transfer to any other provider.
Do credits expire?
It depends on the vendor and most do not say. Cartesia states plainly that they do not: “You will keep the credits in your account until you use them”, and “Existing credits remain: Nothing is lost when you upgrade or downgrade.” Where a vendor is silent, treat the monthly allowance as use-it-or-lose-it until you have it in writing.
Does markup count toward my bill?
With two vendors, yes, and both say so. Google counts every SSML tag except <mark> as billable characters, along with spaces and newlines. Cartesia counts each break tag as one credit.
Are these prices likely to still be current?
Some will not be. Two of the ten publish no effective date on any pricing page, one serves a live payload with no version or history, and one carries the caveat that its prices “are estimates only and are not intended as actual price quotes”. Every figure here is dated to September 1, 2026 for that reason, and this page is re-checked on the site’s standing schedule.
Sources and what we could not verify
Every rate on this page was read on September 1, 2026 from the vendor’s own pricing page, its published documentation, its public price API, or — where the pricing page renders nothing — the vendor’s own JSON pricing endpoints, whose provenance is recorded in the individual reviews. Nothing here is drawn from a third-party price aggregator. Per-provider detail, with source URLs, is in the ten reviews linked throughout, and the verification method is described on the methodology page.
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- Whether any of these prices are stable. Two vendors publish no effective or updated date on any pricing page; Murf’s figures come from a live payload with no version and no dated archive, so we publish them as of the check date and make no claim that they are unchanged from our previous check.
- The Murf unit discrepancy. Murf’s own payload says 10,000 characters per unit while its page code labels the same rates per 1,000. We publish the help-centre figure that two Murf sources agree on and leave the contradiction standing.
- Azure’s “Neural HD Flash” rate. The pricing page groups Flash with Neural, but a case-insensitive search of every field of all 558 items in the Retail Prices API returns zero matches for “flash”. No Microsoft sentence states that Flash is priced identically to Neural, so we do not.
- Resemble AI’s synthesis rates — now resolved, September 2, 2026. No rate is published because the vendor does not sell voice generation. The figures we found in a changelog and a product FAQ are a record of what was once published, not live rates.
- Any cost per minute or per hour of audio, except Cartesia’s own published figure. Converting text to time requires a speaking rate no vendor publishes for billing purposes.
- OpenAI’s newer speech model in characters. gpt-4o-mini-tts is priced in audio tokens with no published conversion.
- Whether any of these rates buy comparable output. No listening test has been run. This is a price comparison and nothing else.
Change log
September 2, 2026 — First publication. All rates read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · This page shows no audio and reports no listening results, because our benchmark has not run. When it does, price will be set against measured output and this page is revisited with a dated entry (methodology). Prices change without notice; corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.