Amazon Polly vs Google Cloud Text-to-Speech 2026
Published September 2, 2026 · This is a documentation-based comparison. Every figure below was read from the two vendors’ own pricing pages, documentation and terms on September 1, 2026 and carries that check date. No listening test has been run, so nothing here compares how the two providers sound.
The short answer
These two list at exactly the same price, so price is not the decision. Amazon Polly’s Standard voices and Google’s Standard and WaveNet voices both cost $4.00 per million characters (Polly pricing, Google pricing, both checked September 1, 2026). Neither has a subscription. Both bill per character, monthly, with no seats and no minimum.
The decision is made by three things instead. Google’s free tier is larger in practice and simpler to reason about. AWS answers the ownership question in a numbered clause where Google’s position is genuinely unresolved on its own pages. And AWS uses your text to improve its models by default, unless an organization administrator turns that off.
Neither will let you clone a voice. Polly’s Brand Voice is a custom engagement with no published price; Google’s Instant Custom Voice is allow-listed behind sales. If cloning is the requirement, this is the wrong pair.
Comparison basis: official documentation only, checked September 1, 2026 — no accounts created, no API calls made, no audio generated, no scores assigned (how we verify).
| Amazon Polly | Google Cloud Text-to-Speech | |
|---|---|---|
| Cheapest rate | $4.00 per 1M characters (Standard) | $4.00 per 1M characters (Standard and WaveNet) |
| Mid tier | $16.00 per 1M (Neural) | $16 per 1M (Neural2, Polyglot) |
| Premium tiers | $30 per 1M (Generative); $100.00 per 1M (Long-Form) | $30 per 1M (Chirp 3: HD); $160 per 1M (Studio) |
| Newest family | Generative, billed per character | Gemini-TTS, billed per token — a different unit entirely |
| Free tier | 5M characters/month on Standard, duration stated two ways; 1M Neural, 500K Long-Form, 100K Generative, each “for the first 12 months” | 4M characters/month on Standard and WaveNet; 1M on the premium families. Gemini-TTS excluded |
| Account requirement | AWS account; Free Plan accounts expire at six months or when credits run out | Billing account required before the first free character |
| Output ownership | Explicit. Service Terms 50.2: “The output that you generate using AI Services is Your Content” | Unresolved. Cloud TTS is listed under Pre-Trained APIs, not Generative AI Services — so the ownership section may or may not reach it |
| Training on your text | On by default, opt-out via an AWS Organizations policy | Not established for this service from the pages we read |
| Voice cloning | Brand Voice — custom engagement, no published price | Instant Custom Voice — allow-listed, $60 per 1M characters |
| Regional breadth | 24 documented regions on Standard, including a China region; Long-Form in one | Regional coverage across the Cloud platform |
| Audio quality | Not compared. Our listening benchmark has not run, so this page assigns no naturalness verdict to either provider. | |
On this page
Winners by category
Every winner below is a documentation-based judgment from published prices, terms and feature documentation. None is a test result, and the categories that require listening have no winner.
| Category | Winner | Why |
|---|---|---|
| Headline price | Tie | Both list at $4.00 per million characters on their cheapest voices |
| Free tier | Google Cloud | 4M characters a month with no duration ambiguity, against a 5M allowance AWS dates two different ways |
| Output ownership | Amazon Polly | A numbered clause says the output is yours; Google’s equivalent section may not even apply to this service |
| Data privacy by default | Google Cloud | AWS uses your text to improve its models unless an administrator opts out |
| Pricing clarity | Amazon Polly | Every engine is billed per character. Google’s newest family is billed per token, which cannot be compared without measuring |
| Premium ceiling cost | Amazon Polly | Its dearest engine is $100 per million against Google’s $160 Studio voices |
| Cheap repeated audio | Amazon Polly | Caching and replay are explicitly free, which changes IVR economics |
| Expressive control | Neither cleanly | Polly’s richest SSML works only on its oldest engine; Google spreads control across families |
| Voice cloning access | Neither | Both require a sales conversation; only Google publishes a price for it |
| Naturalness | No winner | Requires listening. Our benchmark has not run |
| Pronunciation and difficult text | No winner | Requires listening. Our benchmark has not run |
Price: identical headline, different ladders
At the entry level these two are the same number. Amazon prices Standard voices at “$4.00 per 1 million characters for speech or Speech Marks requests”. Google prices Standard and WaveNet at “US$0.000004 per character (US$4 per 1 million characters)”. There is nothing to choose between them on the cheapest tier.
The ladders above that diverge in ways that matter if you need more than the base voices.
- Mid tier is also a tie: Polly’s Neural at $16.00 per million against Google’s Neural2 and Polyglot at $16 per million.
- Premium splits. Polly’s Generative is $30 per million; Google’s Chirp 3: HD is also $30. But Polly’s dearest engine, Long-Form, is $100 per million, while Google’s Studio voices reach $160 — the most expensive per-character rate in our whole comparison set.
- Google adds a second billing unit. Its Gemini-TTS family is priced in tokens rather than characters — for example $0.50 per million input text tokens and $10.00 per million output audio tokens on Gemini 2.5 Flash TTS, with the footnote that “Audio tokens correspond to 25 tokens per second of audio”. You cannot place that on the same axis as a per-character rate without measuring real usage, and we do not attempt it.
One AWS caveat that does not apply to Google in the same form: Polly’s pricing page carries a second, higher rate card for AWS GovCloud (US), at $4.80 for Standard and $19.20 for Neural TTS, with only those two engines listed. A government buyer quoting the commercial rates understates their cost by 20%.
And one Polly advantage worth real money: its pricing page states you “can cache and replay Amazon Polly’s generated speech at no additional cost”. For IVR trees, station announcements and any prompt library, that removes the recurring cost entirely after first generation.
Free tiers, where the real difference is
On paper Amazon offers more: five million characters a month on Standard against Google’s four million. In practice Google’s is the one you can plan around, for two reasons.
AWS states the duration two different ways. Its pricing page attaches “for the first 12 months” explicitly to the Neural, Long-Form and Generative allowances, but not to the Standard one, which reads simply: “For Amazon Polly’s Standard voices, the free tier includes 5 million characters per month for speech or Speech Marks requests.” The Polly FAQ then describes the whole arrangement as time-limited: “Upon sign-up, new Amazon Polly customers can synthesize millions of characters for free each month for the first 12 months.” Whether Standard is perpetual or expires at twelve months therefore has two official answers, and we do not reconcile them.
AWS layers a newer scheme on top. New customers receive up to $200 in Free Tier credits, and the AWS Free Tier Terms state that Free Plan accounts expire “(1) six months from the date you opened your account, or (2) once you have exhausted your Free Tier Credits, whichever comes first”, adding that such accounts “are provided for evaluation purposes and should not be used for processing sensitive data”. How the six-month Free Plan interacts with the twelve-month Polly allowances is not reconciled anywhere we read.
Google’s is simpler. Four million characters a month on Standard and WaveNet, one million on the premium families, recurring, producing ordinary downloadable audio. The condition is procedural rather than temporal: “You must enable billing to use Text-to-Speech, and will be automatically charged if your usage exceeds the number of free characters allowed per month.” A card comes before the first free character. Note that the Gemini-TTS models sit outside the free tier entirely.
What this means for you: if your project needs a predictable recurring free allowance you can budget around for years, Google’s four million is the safer assumption than Amazon’s five. Budget Polly on the twelve-month reading unless you confirm otherwise against your own billing console.
Ownership and data use
This is where the two genuinely part company, and it runs opposite to the free-tier result.
AWS answers the ownership question in a numbered contract clause. Service Terms 50.2, verbatim: “The output that you generate using AI Services is Your Content. Due to the nature of machine learning, output may not be unique across customers and the Services may generate the same or similar results across customers.” The Customer Agreement’s definition of Your Content covers “any computational results” you derive through the Services, which is the language that makes synthesized audio yours.
Google’s position on this service is genuinely unresolved on its own pages. Cloud Text-to-Speech appears in Google’s Services Summary under “Pre-Trained APIs”, not under “Generative AI Services”. That matters because the Service Specific Terms section governing ownership, prohibited use and output retention applies to Generative AI Services — and its own catch-all note extends that heading to “any Generally Available generative AI features of a Service” without saying whether Chirp 3: HD or Gemini-TTS qualify. No Google page we read settles it. The general grant in Cloud Terms §1.1 permits use of the Services; a specific ownership assignment for synthesized audio we did not find.
On training, the direction reverses. AWS Service Terms 50.3 provides for AWS using and storing content processed by its AI Services to develop and improve them, naming Amazon Polly with no tier qualifier. Opting out requires an AWS Organizations policy set by an administrator — there is no toggle in Polly. Note also that AWS’s own Polly product page states it “does not retain the content of your text submissions”, which points the other way; the Service Terms govern. We could not establish an equivalent default-training position for Google Cloud Text-to-Speech from the pages we read, and we do not assert one either way.
What this means for you: if your legal review wants a sentence saying you own the audio, AWS has one and Google does not. If your privacy review objects to inputs improving a vendor’s models, AWS requires you to act at the organization level, and Google’s position for this service is something you should confirm directly rather than infer.
Voice cloning: two locked doors
Neither provider sells cloning self-serve, and they lock the door differently.
Google publishes the price but not the access. Instant Custom Voice carries a documented rate of $60 per million characters, and its documentation states: “Access to Instant Custom Voice is restricted to allow-listed users. To request access, contact a member of the sales team.” You know what it costs before you ask.
AWS publishes neither. Brand Voice is described as “a custom engagement where you work with the Amazon Polly team to build an Neural Text-to-Speech (NTTS) voice for the exclusive use of your organization”. The Polly pricing page contains no occurrence of “Brand Voice” at all, and no minimum recording length, turnaround, commitment or price is published anywhere we read.
If self-serve cloning is what you need, neither of these is the answer. ElevenLabs documents it from $6 a month and Cartesia from $5 (both checked September 1, 2026).
Engineering constraints worth knowing
Several documented limits will shape an implementation more than the price will.
- Polly’s engine parameter is optional and defaults to the oldest engine. “If you don’t provide an engine, the standard engine is selected by default.” With a twenty-five-fold price spread and no IAM condition keys to constrain it, engine selection deserves a code review.
- Polly’s expressive SSML runs backwards. Whisper, emphasis, breaths and vocal-tract-length tags are documented as available on the cheapest engine and “Not available” on neural, long-form and generative.
- Polly’s generative engine produces no speech marks at all, ruling it out for lip-sync or word-level highlighting, and its Long-Form engine is documented in exactly one AWS region.
- Neither publishes a latency figure for text-to-speech, and we found no service-level agreement for Polly on any page we read.
- Google’s documentation host moved. Older citations pointing at
cloud.google.com/text-to-speech/docs/…now redirect todocs.cloud.google.com. The pricing page and terms pages are unaffected.
Voice quality, pronunciation and naturalness
We have not compared how these two providers sound, and we will not imply otherwise.
Our standardized listening benchmark is designed but not running. Until it does, this page carries no naturalness verdict, no pronunciation comparison and no scores for either provider. Both vendors make quality claims of their own — AWS describes its neural engine as capable of “even higher quality voices than its standard voices” — and those belong to them.
This is the honest limit of a documentation-based comparison between two providers whose headline prices are identical. Everything that can be settled by reading is above. The thing that would actually separate them for most buyers — how the voices sound at $4 a million — cannot be.
Choose Polly if · Choose Google if
Choose Amazon Polly if:
- Your legal review needs ownership in a numbered clause. Service Terms 50.2 states it; Google’s equivalent may not even apply to this service.
- You generate the same audio repeatedly. Caching and replay are explicitly free, which is decisive for IVR and prompt libraries.
- You are already inside AWS, or need GovCloud or a China region.
- Your premium ceiling matters. $100 per million on Long-Form beats Google’s $160 Studio rate.
Choose Google Cloud Text-to-Speech if:
- You want a free allowance you can plan around. Four million characters a month, recurring, with no duration ambiguity.
- You object to inputs improving a vendor’s models by default, and cannot rely on an organization-level opt-out being set.
- You are already on Google Cloud and want its IAM and regional model for free in engineering terms.
- You want cloning priced transparently even if gated. $60 per million is published; AWS publishes nothing for Brand Voice.
Sources and what we could not verify
Every figure on this page was read from the two vendors’ official pages on September 1, 2026. Full detail is in the individual reviews: Amazon Polly review and Google Cloud Text-to-Speech review. How we verify anything is described on the methodology page.
| Source | Used for |
|---|---|
| Amazon Polly pricing | Engine rates, GovCloud rates, free-tier allowances, caching and replay |
| Amazon Polly FAQs | The conflicting free-tier duration wording |
| AWS Service Terms | Clauses 50.2 and 50.3 — ownership and default data use |
| AWS Free Tier Terms | Free Plan expiry and the evaluation-purposes language |
| Polly SSML support | Per-engine availability of expressive tags |
| Amazon Polly features | The Brand Voice description |
| Google Cloud TTS pricing | Per-character and per-token rates, free-tier bands, billing-account requirement |
| Google Cloud Platform Terms | The §1.1 services grant |
| Google Services Summary | Cloud TTS classified under Pre-Trained APIs (last modified August 27, 2026) |
| Instant Custom Voice docs | The sales allow-list restriction on cloning |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- How either provider sounds. No listening test has been run.
- Whether Amazon’s Standard free tier is perpetual or twelve months. Two official pages disagree; both published.
- How AWS’s six-month Free Plan interacts with Polly’s twelve-month allowances. No page reconciles them.
- Whether Google’s Generative AI ownership terms reach Cloud Text-to-Speech. It is listed under Pre-Trained APIs, and the catch-all extending that heading is not resolved on any Google page.
- Google’s default position on using inputs for model improvement in this service. Not established from the pages we read; we assert nothing either way.
- Brand Voice pricing, sample requirements or turnaround. Not published anywhere.
- Any per-character comparison against Google’s Gemini-TTS models. They bill in tokens and no conversion is published.
- Any latency figure or SLA for Polly. Not published on any page we read.
Change log
September 2, 2026 — First publication as a documentation-based comparison. All figures read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · When our audio benchmark runs, this page gains a head-to-head listening test using the same prompts and documented conditions for both providers, and the naturalness and pronunciation rows gain real winners (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.