Amazon Polly vs Google Cloud Text-to-Speech 2026

✓ Docs-verified · Sep 1, 2026◌ Not audio-tested△ What we could not verify

Published September 2, 2026 · This is a documentation-based comparison. Every figure below was read from the two vendors’ own pricing pages, documentation and terms on September 1, 2026 and carries that check date. No listening test has been run, so nothing here compares how the two providers sound.

The short answer

These two list at exactly the same price, so price is not the decision. Amazon Polly’s Standard voices and Google’s Standard and WaveNet voices both cost $4.00 per million characters (Polly pricing, Google pricing, both checked September 1, 2026). Neither has a subscription. Both bill per character, monthly, with no seats and no minimum.

The decision is made by three things instead. Google’s free tier is larger in practice and simpler to reason about. AWS answers the ownership question in a numbered clause where Google’s position is genuinely unresolved on its own pages. And AWS uses your text to improve its models by default, unless an organization administrator turns that off.

Neither will let you clone a voice. Polly’s Brand Voice is a custom engagement with no published price; Google’s Instant Custom Voice is allow-listed behind sales. If cloning is the requirement, this is the wrong pair.

Comparison basis: official documentation only, checked September 1, 2026 — no accounts created, no API calls made, no audio generated, no scores assigned (how we verify).

Head-to-head on documented facts — verified September 1, 2026
Amazon PollyGoogle Cloud Text-to-Speech
Cheapest rate $4.00 per 1M characters (Standard) $4.00 per 1M characters (Standard and WaveNet)
Mid tier $16.00 per 1M (Neural) $16 per 1M (Neural2, Polyglot)
Premium tiers $30 per 1M (Generative); $100.00 per 1M (Long-Form) $30 per 1M (Chirp 3: HD); $160 per 1M (Studio)
Newest family Generative, billed per character Gemini-TTS, billed per token — a different unit entirely
Free tier 5M characters/month on Standard, duration stated two ways; 1M Neural, 500K Long-Form, 100K Generative, each “for the first 12 months” 4M characters/month on Standard and WaveNet; 1M on the premium families. Gemini-TTS excluded
Account requirement AWS account; Free Plan accounts expire at six months or when credits run out Billing account required before the first free character
Output ownership Explicit. Service Terms 50.2: “The output that you generate using AI Services is Your Content” Unresolved. Cloud TTS is listed under Pre-Trained APIs, not Generative AI Services — so the ownership section may or may not reach it
Training on your text On by default, opt-out via an AWS Organizations policy Not established for this service from the pages we read
Voice cloning Brand Voice — custom engagement, no published price Instant Custom Voice — allow-listed, $60 per 1M characters
Regional breadth 24 documented regions on Standard, including a China region; Long-Form in one Regional coverage across the Cloud platform
Audio quality Not compared. Our listening benchmark has not run, so this page assigns no naturalness verdict to either provider.
On this page

Winners by category

Every winner below is a documentation-based judgment from published prices, terms and feature documentation. None is a test result, and the categories that require listening have no winner.

Category winners — documentation-based judgments, verified September 1, 2026
CategoryWinnerWhy
Headline priceTieBoth list at $4.00 per million characters on their cheapest voices
Free tierGoogle Cloud4M characters a month with no duration ambiguity, against a 5M allowance AWS dates two different ways
Output ownershipAmazon PollyA numbered clause says the output is yours; Google’s equivalent section may not even apply to this service
Data privacy by defaultGoogle CloudAWS uses your text to improve its models unless an administrator opts out
Pricing clarityAmazon PollyEvery engine is billed per character. Google’s newest family is billed per token, which cannot be compared without measuring
Premium ceiling costAmazon PollyIts dearest engine is $100 per million against Google’s $160 Studio voices
Cheap repeated audioAmazon PollyCaching and replay are explicitly free, which changes IVR economics
Expressive controlNeither cleanlyPolly’s richest SSML works only on its oldest engine; Google spreads control across families
Voice cloning accessNeitherBoth require a sales conversation; only Google publishes a price for it
NaturalnessNo winnerRequires listening. Our benchmark has not run
Pronunciation and difficult textNo winnerRequires listening. Our benchmark has not run

Price: identical headline, different ladders

At the entry level these two are the same number. Amazon prices Standard voices at “$4.00 per 1 million characters for speech or Speech Marks requests”. Google prices Standard and WaveNet at “US$0.000004 per character (US$4 per 1 million characters)”. There is nothing to choose between them on the cheapest tier.

The ladders above that diverge in ways that matter if you need more than the base voices.

One AWS caveat that does not apply to Google in the same form: Polly’s pricing page carries a second, higher rate card for AWS GovCloud (US), at $4.80 for Standard and $19.20 for Neural TTS, with only those two engines listed. A government buyer quoting the commercial rates understates their cost by 20%.

And one Polly advantage worth real money: its pricing page states you “can cache and replay Amazon Polly’s generated speech at no additional cost”. For IVR trees, station announcements and any prompt library, that removes the recurring cost entirely after first generation.

Free tiers, where the real difference is

On paper Amazon offers more: five million characters a month on Standard against Google’s four million. In practice Google’s is the one you can plan around, for two reasons.

AWS states the duration two different ways. Its pricing page attaches “for the first 12 months” explicitly to the Neural, Long-Form and Generative allowances, but not to the Standard one, which reads simply: “For Amazon Polly’s Standard voices, the free tier includes 5 million characters per month for speech or Speech Marks requests.” The Polly FAQ then describes the whole arrangement as time-limited: “Upon sign-up, new Amazon Polly customers can synthesize millions of characters for free each month for the first 12 months.” Whether Standard is perpetual or expires at twelve months therefore has two official answers, and we do not reconcile them.

AWS layers a newer scheme on top. New customers receive up to $200 in Free Tier credits, and the AWS Free Tier Terms state that Free Plan accounts expire “(1) six months from the date you opened your account, or (2) once you have exhausted your Free Tier Credits, whichever comes first”, adding that such accounts “are provided for evaluation purposes and should not be used for processing sensitive data”. How the six-month Free Plan interacts with the twelve-month Polly allowances is not reconciled anywhere we read.

Google’s is simpler. Four million characters a month on Standard and WaveNet, one million on the premium families, recurring, producing ordinary downloadable audio. The condition is procedural rather than temporal: “You must enable billing to use Text-to-Speech, and will be automatically charged if your usage exceeds the number of free characters allowed per month.” A card comes before the first free character. Note that the Gemini-TTS models sit outside the free tier entirely.

What this means for you: if your project needs a predictable recurring free allowance you can budget around for years, Google’s four million is the safer assumption than Amazon’s five. Budget Polly on the twelve-month reading unless you confirm otherwise against your own billing console.

Ownership and data use

This is where the two genuinely part company, and it runs opposite to the free-tier result.

AWS answers the ownership question in a numbered contract clause. Service Terms 50.2, verbatim: “The output that you generate using AI Services is Your Content. Due to the nature of machine learning, output may not be unique across customers and the Services may generate the same or similar results across customers.” The Customer Agreement’s definition of Your Content covers “any computational results” you derive through the Services, which is the language that makes synthesized audio yours.

Google’s position on this service is genuinely unresolved on its own pages. Cloud Text-to-Speech appears in Google’s Services Summary under “Pre-Trained APIs”, not under “Generative AI Services”. That matters because the Service Specific Terms section governing ownership, prohibited use and output retention applies to Generative AI Services — and its own catch-all note extends that heading to “any Generally Available generative AI features of a Service” without saying whether Chirp 3: HD or Gemini-TTS qualify. No Google page we read settles it. The general grant in Cloud Terms §1.1 permits use of the Services; a specific ownership assignment for synthesized audio we did not find.

On training, the direction reverses. AWS Service Terms 50.3 provides for AWS using and storing content processed by its AI Services to develop and improve them, naming Amazon Polly with no tier qualifier. Opting out requires an AWS Organizations policy set by an administrator — there is no toggle in Polly. Note also that AWS’s own Polly product page states it “does not retain the content of your text submissions”, which points the other way; the Service Terms govern. We could not establish an equivalent default-training position for Google Cloud Text-to-Speech from the pages we read, and we do not assert one either way.

What this means for you: if your legal review wants a sentence saying you own the audio, AWS has one and Google does not. If your privacy review objects to inputs improving a vendor’s models, AWS requires you to act at the organization level, and Google’s position for this service is something you should confirm directly rather than infer.

Voice cloning: two locked doors

Neither provider sells cloning self-serve, and they lock the door differently.

Google publishes the price but not the access. Instant Custom Voice carries a documented rate of $60 per million characters, and its documentation states: “Access to Instant Custom Voice is restricted to allow-listed users. To request access, contact a member of the sales team.” You know what it costs before you ask.

AWS publishes neither. Brand Voice is described as “a custom engagement where you work with the Amazon Polly team to build an Neural Text-to-Speech (NTTS) voice for the exclusive use of your organization”. The Polly pricing page contains no occurrence of “Brand Voice” at all, and no minimum recording length, turnaround, commitment or price is published anywhere we read.

If self-serve cloning is what you need, neither of these is the answer. ElevenLabs documents it from $6 a month and Cartesia from $5 (both checked September 1, 2026).

Engineering constraints worth knowing

Several documented limits will shape an implementation more than the price will.

Voice quality, pronunciation and naturalness

We have not compared how these two providers sound, and we will not imply otherwise.

Our standardized listening benchmark is designed but not running. Until it does, this page carries no naturalness verdict, no pronunciation comparison and no scores for either provider. Both vendors make quality claims of their own — AWS describes its neural engine as capable of “even higher quality voices than its standard voices” — and those belong to them.

This is the honest limit of a documentation-based comparison between two providers whose headline prices are identical. Everything that can be settled by reading is above. The thing that would actually separate them for most buyers — how the voices sound at $4 a million — cannot be.

Choose Polly if · Choose Google if

Choose Amazon Polly if:

Choose Google Cloud Text-to-Speech if:

Sources and what we could not verify

Every figure on this page was read from the two vendors’ official pages on September 1, 2026. Full detail is in the individual reviews: Amazon Polly review and Google Cloud Text-to-Speech review. How we verify anything is described on the methodology page.

Official sources consulted — all checked September 1, 2026
SourceUsed for
Amazon Polly pricingEngine rates, GovCloud rates, free-tier allowances, caching and replay
Amazon Polly FAQsThe conflicting free-tier duration wording
AWS Service TermsClauses 50.2 and 50.3 — ownership and default data use
AWS Free Tier TermsFree Plan expiry and the evaluation-purposes language
Polly SSML supportPer-engine availability of expressive tags
Amazon Polly featuresThe Brand Voice description
Google Cloud TTS pricingPer-character and per-token rates, free-tier bands, billing-account requirement
Google Cloud Platform TermsThe §1.1 services grant
Google Services SummaryCloud TTS classified under Pre-Trained APIs (last modified August 27, 2026)
Instant Custom Voice docsThe sales allow-list restriction on cloning

What we could not verify

Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:

Change log

September 2, 2026 — First publication as a documentation-based comparison. All figures read from official vendor sources on September 1, 2026 and dated accordingly.

Published September 2, 2026 · When our audio benchmark runs, this page gains a head-to-head listening test using the same prompts and documented conditions for both providers, and the naturalness and pronunciation rows gain real winners (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.