Amazon Polly Review 2026
Published September 1, 2026 · Every fact below was read from AWS’s own pricing page, developer guide, Customer Agreement and Service Terms on September 1, 2026, with the source linked beside the claim.
Quick verdict
Amazon Polly is the cheapest documented way to synthesize speech at volume, and the clearest documented answer to “who owns the audio” — provided you treat it as infrastructure rather than a product. Standard voices list at $4.00 per 1 million characters, and AWS states plainly that output you generate is yours (pricing; Service Terms 50.2, both checked September 1, 2026). There is no editor, no self-serve voice cloning, and no subscription — you get an API, an IAM policy and a bill.
The dated warning that matters more: your text is used to improve AWS models by default, on the free tier and paid alike, unless an administrator sets an AWS Organizations opt-out policy. Service Terms 50.3 names Amazon Polly with no tier qualifier while carving out tiers for other products in the same clause (Service Terms, checked September 1, 2026). AWS’s own product page says something different — see commercial usage, licensing, privacy.
Verdict basis: official documentation only, checked September 1, 2026 — no account, no API call, no listening test, no numeric score (how we verify).
| Best for | Developers already inside AWS who need high-volume synthesis at the lowest documented list price, with output ownership stated in the contract |
|---|---|
| Not best for | Anyone who wants an editor, self-serve voice cloning, a published latency figure, or expressive SSML on a current-generation engine |
| Starting price | No subscription. Pay-as-you-go from $4.00 per 1 million characters on Standard voices; $16.00 Neural; $30 Generative; $100.00 Long-Form (pricing) |
| Commercial use | Not addressed as a permission or a prohibition anywhere AWS documents it. The word “commercial” does not appear on the pricing page or the FAQ. Rights rest on the ownership clause instead (Service Terms 50.2) |
| Voice cloning | No self-serve cloning at any price. “Brand Voice” is a custom engagement with the Polly team; no price is published anywhere (features) |
| Test status | Documentation-verified. Our listening benchmark has not run, so there is no independent voice-quality score and no measured latency on this page. |
Every price, licensing term and feature claim below was read from AWS’s official pages on September 1, 2026, with the source linked beside the fact. Our standardized audio benchmark is designed but not yet running, so this review contains no listening-test claims, no naturalness verdicts and no scores. Where AWS makes a quality claim, it is quoted and attributed to AWS.
Strongest advantages:
- The lowest documented list price in our coverage: $4.00 per 1 million characters on Standard voices (pricing, checked September 1, 2026).
- An explicit ownership clause in a numbered contract term, not a marketing FAQ: “The output that you generate using AI Services is Your Content” (Service Terms 50.2).
- Caching and replay are permitted at no extra charge, stated on the pricing page itself — which matters for IVR prompts and any repeated audio (pricing).
- The broadest documented regional footprint of any engine here: Standard voices list 24 regions, including one in China (standard voices).
- A recurring monthly free allowance rather than a one-off trial: 5 million characters per month on Standard voices (pricing) — though its duration is stated two ways, see below.
Important disadvantages:
- Your text trains AWS models by default unless an administrator opts out through AWS Organizations (Service Terms 50.3).
- No self-serve voice cloning exists at any price; Polly’s ten API operations contain no create-voice call (API operations).
- Expressive SSML — whispering, emphasis, breaths, vocal-tract length — is documented as available only on the oldest and cheapest engine (supported SSML tags).
- The most expensive engine, Long-Form at $100.00 per 1 million characters, is documented in exactly one AWS region (long-form voices).
- AWS publishes no latency figure and no service-level agreement for Polly on any page we read.
- AWS documents that voices may change over time, which matters for serialized content (generative voices).
On this page
Who should use it — and who should look elsewhere
Choose Amazon Polly if:
- You are already an AWS customer and volume is your problem. Standard voices at $4.00 per 1 million characters are the lowest documented list price in our coverage, and there is no plan to outgrow — billing is per character, monthly, with no seats and no minimum (pricing, checked September 1, 2026).
- Your legal review needs ownership in a numbered clause. Service Terms 50.2 states it directly: “The output that you generate using AI Services is Your Content” (Service Terms, last updated August 20, 2026, checked September 1, 2026). Several competitors answer this question only in undated marketing copy.
- You generate the same audio repeatedly. The pricing page states you “can cache and replay Amazon Polly’s generated speech at no additional cost”, and the FAQ confirms static prompts replayed many times carry no extra charge (pricing, FAQ). For IVR and navigation prompts this changes the economics entirely.
- You need synthesis inside a specific region, including GovCloud or China. Standard voices are documented in 24 regions, the broadest footprint of any engine on the service (standard voices). Note that GovCloud pricing is 20% higher and covers only two engines — see pricing.
Look elsewhere if:
- You want to clone a voice without a sales conversation. Polly has no self-serve cloning at any tier. If cloning is the requirement, ElevenLabs documents instant cloning from its $6/month Starter plan (elevenlabs.io/pricing, checked September 1, 2026).
- You need an editor rather than an API. Polly ships no timeline, no project view and no multi-voice script tool. Murf AI sells exactly that, from $29/month (murf.ai/pricing, checked September 1, 2026).
- You cannot accept your inputs being used for model improvement, and you have no AWS Organizations administrator. The opt-out is an organization-level policy, not a checkbox in the Polly console (Service Terms 50.3).
- Your buying decision needs a published latency number or an SLA. AWS publishes neither for Polly on any page we read. If a documented figure is contractually necessary, this is not the service to build that requirement on.
Current price, free plan, commercial rights
The short version: there is no subscription; you pay per character at one of four engine rates, the cheapest being $4.00 and the most expensive $100.00 per 1 million characters; there is a recurring monthly free allowance whose duration AWS states inconsistently; and no AWS page grants or refuses commercial use in those words — your rights come from the ownership clause instead.
| Engine | Price per 1M characters | Free allowance, as AWS words it | Notable limit |
|---|---|---|---|
| Standard | $4.00 | “the free tier includes 5 million characters per month for speech or Speech Marks requests” — no duration stated | The API default if you omit the engine parameter |
| Neural | $16.00 | “1 million characters per month … for the first 12 months” | Expressive SSML tags documented as not available |
| Generative | $30 | “100 thousand characters per month for speech requests, for the first 12 months” | No speech marks at all; the only engine with bidirectional streaming |
| Long-Form | $100.00 | “500 thousand characters per month … for the first 12 months” | Documented in one region only: US East (N. Virginia) |
Prices are quoted verbatim from the pay-as-you-go block on the Polly pricing page, checked September 1, 2026. AWS writes the Generative rate as “$30” in the prose and “$30.00” in its own examples table; we reproduce both rather than pick one.
A second price list applies to government customers. The same pricing page carries an “AWS GovCloud (US) Pricing Details” section with higher rates and fewer engines: Standard voices at $4.80 and Neural TTS voices at $19.20 per 1 million characters, with no Long-Form or Generative pricing given. Quoting the four commercial rates at a GovCloud buyer understates their cost by 20%.
What this means for you: the engine parameter is the single largest cost decision on this service, and it is easy to get wrong in both directions. Omit it and AWS gives you Standard by default — verbatim, “If you don’t provide an engine, the standard engine is selected by default” (SynthesizeSpeech reference). Set it to long-form without checking, and you are paying twenty-five times the Standard rate for the same character count.
The free tier: AWS states its duration two different ways
This is a genuine contradiction between two official AWS pages, and we record it rather than resolve it.
The pricing page attaches the phrase “for the first 12 months” explicitly to the Neural, Long-Form and Generative allowances. It does not attach that phrase to the Standard allowance, which reads simply: “For Amazon Polly’s Standard voices, the free tier includes 5 million characters per month for speech or Speech Marks requests.”
The Polly FAQ describes the whole arrangement as time-limited: “Upon sign-up, new Amazon Polly customers can synthesize millions of characters for free each month for the first 12 months.”
Whether the Standard free allowance is perpetual or expires at twelve months therefore has two official answers. We could not settle it from the rendered text of either page, and we do not average or pick the more favourable one. Treat the 12-month reading as the safe assumption for budgeting, and verify against your own account’s billing console before you plan around a perpetual free allowance.
A separate and newer complication: AWS now describes a credit-based free tier alongside the character allowances. The pricing page and FAQ both state that “Starting July 15, 2025, new AWS customers will receive up to $200 in AWS Free Tier credits, which can be applied towards eligible AWS services, including Amazon Polly.” The AWS Free Tier Terms add two conditions that a Polly buyer should read before relying on any of it: Free Plan accounts expire “(1) six months from the date you opened your account, or (2) once you have exhausted your Free Tier Credits, whichever comes first”, and “Free Plan accounts are provided for evaluation purposes and should not be used for processing sensitive data.” How the six-month Free Plan interacts with the twelve-month Polly allowances is not reconciled on any page we read.
Commercial rights: an ownership clause, not a permission
Most vendors in this market answer the commercial-use question directly, in a licence clause that says yes or no. AWS does not. The word “commercial” appears zero times on both the Polly pricing page and the Polly FAQ, checked September 1, 2026. There is no clause granting commercial use and none prohibiting it.
What exists instead is an ownership statement. AWS Service Terms 50.2 reads, verbatim:
“The output that you generate using AI Services is Your Content. Due to the nature of machine learning, output may not be unique across customers and the Services may generate the same or similar results across customers.”
The AWS Customer Agreement defines “Your Content” to include “any computational results that you or any End User derive from the foregoing through their use of the Services” — which is the language that makes synthesized audio yours (Customer Agreement, Section 12, last updated August 14, 2026). Clause 6.1 adds that AWS “obtain[s] no rights under this Agreement from you (or your licensors) to Your Content.”
Read the second sentence of 50.2 as carefully as the first. AWS is telling you that the same input may produce the same output for another customer. You own what you generate; you are not being promised it is unique to you.
What this review covers
This review covers the four engines Amazon Polly sells today, named by their API identifiers as the request parameter enumerates them: standard, neural, long-form and generative (SynthesizeSpeech reference, checked September 1, 2026). It covers the published pricing for commercial regions and for GovCloud (US), the free-tier allowances and their stated durations, the ownership and data-use terms that govern output, the documented feature differences between engines, and the API surface, quotas and regional availability.
It does not evaluate how any Polly voice sounds. We have generated no audio, created no AWS account and made no API call for this review. Rows a review of ours normally carries once our benchmark runs — Tested, Model tested, Voice tested, Benchmark cost — are absent here rather than filled with placeholders, because that evidence does not exist yet.
One structural fact governs every legal citation on this page: Amazon Polly has no dedicated numbered section in the AWS Service Terms. A full-text search of that document returns “Amazon Polly” exactly twice, in clauses 50.1 and 50.3. Polly is governed as one of the enumerated “AI Services” under section 50. The Polly FAQ’s answer about children’s privacy refers to “the Amazon Polly Service Terms” as though a dedicated section exists; it does not.
Benchmark audio status
Samples are pending. Our standardized audio benchmark is designed but not running, so this page publishes no audio, no player and no placeholder where a player would go (methodology). What we verified instead is documentation: engine availability by region, output formats and sample rates, SSML tag support by engine, quota and concurrency ceilings, and every price and licence term cited above, all read on September 1, 2026.
AWS publishes no measured latency figure for Polly, so this review contains no speed claim of any kind — not even a vendor one to quote. Any cost figure here is a documented list price, not a measurement of what a real workload cost us.
Assessment by dimension
No numeric scores appear below, because score bands are only meaningful against real measurements and a precise-looking number produced from documentation would be false precision. The judgments are qualitative and drawn from what AWS documents; vendor claims are labelled as such in the cell.
| Dimension | What the documentation supports |
|---|---|
| Price | Strong. $4.00 per 1 million characters on Standard is the lowest documented list price in our coverage; caching and replay carry no additional charge. |
| Output ownership | Strong. Stated in a numbered contract clause (Service Terms 50.2), not in marketing copy. Weakened only by the same clause’s non-uniqueness warning. |
| Data privacy | Weak by default. Inputs may be used to develop and improve AWS services unless an organization-level opt-out is set; AWS’s own product page states the opposite of its Service Terms (see licensing). |
| Voice cloning | Absent self-serve. No API operation creates a voice. Brand Voice exists only as a custom engagement with no published price, sample requirement or turnaround. |
| Expressive control | Inverted. Whisper, emphasis, breaths and vocal-tract-length tags are documented as available on Standard only, and “Not available” on neural, long-form and generative. |
| Regional coverage | Strong on Standard, narrow above it. 24 documented regions for Standard; Long-Form is documented in one. |
| Developer surface | Strong, with AWS assumptions. IAM and SigV4 rather than an API key; ten operations; documented quotas. No service-specific IAM condition keys, so you cannot gate spend by engine through policy. |
| Voice quality | Not assessed. AWS calls its neural engine capable of “even higher quality voices than its standard voices”; that is AWS’s claim, quoted, not our finding. |
| Latency | Not assessed and not published. No millisecond figure appears on any AWS page we read. |
The four engines, and which one you actually get
Polly’s engine parameter is the whole product. It sets the price, the available voices, the SSML tags you may use, the regions you can call from, and whether you get speech marks at all. The canonical enumeration is the request parameter itself, verbatim: “Valid Values: standard | neural | long-form | generative” (SynthesizeSpeech reference, checked September 1, 2026).
Standard is concatenative synthesis. AWS describes it verbatim as an engine that “concatenates phonemes of recorded speech, producing very natural-sounding synthesized speech” — that phrasing is AWS’s, and the grammatical slip is in the original. AWS states an inventory of “40 female and 20 male standard voices in 29 language and language variants” (standard voices). Worth noting for anyone assessing how actively maintained this engine is: the feature list on that same page is published under a header reading “The Amazon Polly standard engine supports the following features (TBD):” — the literal string “(TBD)” is in AWS’s own documentation.
Neural is the long-standing default upgrade. AWS’s claim for it, quoted: an engine “that can produce even higher quality voices than its standard voices.” We do not evaluate that claim.
Generative is the newest engine and the only one supporting bidirectional streaming. It is also the only one AWS documents as capable of producing content you did not ask for. Verbatim, on the safeguard against runaway generation: “This safety feature reduces but does not eliminate the risk. In some edge cases, the model might continue generating random content beyond the stop threshold. Do not assume complete protection” (generative voices). If you are synthesizing unattended at scale, that sentence is the one to take to your risk review.
Long-Form is the premium tier at $100.00 per 1 million characters, and it is documented in exactly one region. Verbatim: “US East (N. Virginia): us-east-1 + Other regions not available” (long-form voices). For a buyer with EU or APAC data-residency requirements, that single line disqualifies the engine regardless of price.
Two engine facts deserve emphasis because they are easy to miss and expensive to discover late. The generative engine produces no speech marks: “Support for generating speech marks is currently not available”, repeated for streaming as “Speech marks are not supported by this operation.” If you are building lip-sync, visemes or word-level highlighting, the newest engine cannot serve you. And AWS documents that voices are not fixed: “any updates to the training data and the model could result in slight variations to the way the voices sound”, published on both the generative and long-form pages. For a serialized audiobook or a multi-year IVR estate, that is a planning constraint, not a footnote.
Voice quality and naturalness
We will not tell you how Amazon Polly sounds. No listening test has been run for this review, and describing synthesized speech from documentation would be inventing evidence.
What documentation does settle is the quality ceiling you can buy, which is a different and more checkable question. AWS documents sample rates in hertz, and does so inconsistently: four official pages give four different sets of valid rates for MP3 and OGG output. The SynthesizeSpeech reference lists 8000, 16000, 22050, 24000, 44100 and 48000; other pages differ. We record the disagreement rather than pick a set. Separately, AWS publishes no MP3 bitrate figure anywhere we read, only sample rates — so a bitrate comparison against providers that do publish one is not possible from official sources.
AWS’s own quality language is a claim, and we quote it as one: the neural engine “can produce even higher quality voices than its standard voices.” Whether that difference matters for your script is exactly what our benchmark exists to answer, and it has not run.
Expression, pacing, control
Polly’s expressive controls run in the opposite direction to the rest of this market: the richest tags work on the cheapest, oldest engine, and are documented as unavailable on the three newer ones.
The SSML tags <emphasis>, <amazon:auto-breaths>, <amazon:effect name="whispered">, <amazon:effect phonation="soft">, <amazon:effect vocal-tract-length> and <prosody amazon:max-duration> are all listed as “Not available” for the neural, long-form and generative engines (supported SSML tags, checked September 1, 2026).
What this means for you: if whispering or fine-grained emphasis is central to your script, the engine that supports it is the concatenative one at $4.00 per 1 million characters — and you should confirm the result meets your bar before committing, since we have not listened to it. If you need a current-generation engine, you are buying a plainer control surface.
Voice cloning
There is no self-serve voice cloning on Amazon Polly at any price. Polly exposes ten API operations — DeleteLexicon, DescribeVoices, GetLexicon, GetSpeechSynthesisTask, ListLexicons, ListSpeechSynthesisTasks, PutLexicon, StartSpeechSynthesisStream, StartSpeechSynthesisTask and SynthesizeSpeech — and none of them creates, trains or uploads a voice.
What AWS offers instead is Brand Voice, described verbatim on the features page as “a custom engagement where you work with the Amazon Polly team to build an Neural Text-to-Speech (NTTS) voice for the exclusive use of your organization… We work with you throughout the entire process to identify the persona, identify an actor or actress and record their speech, and ultimately build and train a model to produce the voice. The voice is then made available to your AWS account ID(s)” (features, checked September 1, 2026).
Everything a buyer would need to evaluate Brand Voice is unpublished. The Polly pricing page contains no occurrence of “Brand Voice” at all. No minimum recording length, script length, turnaround time, minimum commitment or contract term appears on any official page we read, and the FAQ directs the question to an AWS Account Manager. There is also no Brand-Voice-specific consent document: the only consent language AWS publishes is a general prohibition in the Responsible AI Policy, which forbids using the AI/ML services “to depict a person’s voice or likeness without their consent or other appropriate rights, including unauthorized impersonation and non-consensual sexual imagery”. That bullet is unnumbered on the page.
For a consent regime you can actually read before signing, providers that publish their cloning requirements openly are the better comparison — see alternatives.
Languages and voices
AWS states an inventory of “40 female and 20 male standard voices in 29 language and language variants” on the standard-voices page (checked September 1, 2026). We report that as AWS’s count. We could not reconcile it against AWS’s own tables: the page enumerates fewer named standard voices than sixty.
Two further inconsistencies sit inside AWS’s voice documentation, and we record both without resolving either. The Turkish voice Filiz is labelled Male on the standard-voices page and Female on the master voice list. And on the long-form page, the prose names “Daniel, Gregory, and Ruth” while the table immediately beneath it lists Danielle, Gregory, Ruth and Patrick. If you are scripting voice selection programmatically, read the voice list endpoint rather than the prose.
API, SDKs, quotas
Authentication is AWS IAM with SigV4 request signing. There is no Polly API key to paste into a client. That is convenient if you are already inside AWS and an obstacle if you are not.
One consequence is worth stating for anyone planning cost controls: AWS documents no service-specific IAM condition keys for Polly — the service-authorization reference records “Policy condition keys (service-specific): No”. You therefore cannot write an IAM policy that permits Standard but denies Long-Form, or that caps spend by engine. Cost control has to happen elsewhere in your stack.
Documented quota and concurrency ceilings that shape real pipelines, all checked September 1, 2026:
- Batch throughput defaults to 1 transaction per second on the two premium engines. StartSpeechSynthesisTask documents “Generative voice: 1 tps” and “Long-form voice: 1 tps”, against 10 tps for neural and standard. It is adjustable through Service Quotas, but it is the default bottleneck on precisely the engines a bulk audiobook pipeline would choose.
- Bidirectional streaming is generative-only and tightly bounded. Verbatim: “Currently, only the generative engine is supported”; maximum stream duration 10 minutes; eight concurrent requests; and an idle timeout of five seconds between consecutive events, after which the stream closes.
- The engine parameter is optional and defaults to the oldest engine. Omitting it silently selects
standard.
Generation speed and latency
AWS publishes no latency figure for Amazon Polly. There is no millisecond number, no time-to-first-byte, and no p50 or p99 on the pricing page, the product page or the developer guide pages we read on September 1, 2026. There is also no published service-level agreement for the service.
The only quantitative-looking string we found — “50 milliseconds” on the quotas page — sits inside a worked example about request throttling, not a performance claim, and we will not repurpose it as one. AWS’s own speed language is qualitative marketing phrasing about responding in near-real time.
We therefore have nothing to report here, not even a vendor claim to quote. This section will carry measured figures when our benchmark runs.
Pricing and normalized cost
What you pay, and what you get
There is no subscription to normalize. AWS bills monthly for characters processed, with no seats, no annual commitment and no volume discount published on the pricing page — independently confirmed by searching the rendered page, where “seat”, “annual” and “subscription” each occur zero times.
Effective cost
Because billing is purely per character, the effective rate is the list rate: cost = (characters ÷ 1,000,000) × engine rate. A 100,000-character audiobook chapter costs $0.40 on Standard, $1.60 on Neural, $3.00 on Generative and $10.00 on Long-Form, at the rates published on September 1, 2026. Those are arithmetic from list prices, not measurements of a workload we ran.
The important limitation: one parameter moves the bill 25×
The gap between the cheapest and dearest engine is twenty-five to one, and the selection between them is a single optional string in the request. Nothing in AWS’s IAM model lets you constrain it. That combination — a huge price spread, a silent default, and no policy-level guardrail — is the structural risk of running Polly at scale, and it is why engine selection deserves a code review rather than a config default.
Who gets the best value
- High-volume Standard-engine workloads. Nothing else in our coverage lists at $4.00 per 1 million characters.
- Anything cached and replayed — IVR prompts, station announcements, navigation phrases — because replay is free by AWS’s own statement.
- Poor value for occasional creative work. Without an editor, a free usable tier for non-developers, or cloning, a creator is paying in engineering time what they would save in list price.
Commercial usage, licensing, privacy
Who owns what you generate
Quoted above and worth restating exactly: “The output that you generate using AI Services is Your Content” (Service Terms 50.2). The Customer Agreement’s definition of Your Content covers “any computational results” you derive through the Services, and clause 6.1 confirms AWS “obtain[s] no rights under this Agreement” to it.
Two limits sit beside that grant. The same clause 50.2 warns output “may not be unique across customers”. And Polly is not in the list of generative AI services AWS indemnifies under clause 50.10 — we verified its absence from that list by full-text search. AWS makes no affirmative statement that Polly output is uncovered; it simply is not named.
Your text trains AWS models by default
This is the most consequential term on the page for a privacy review, and it is stated three different ways across AWS’s own properties.
Service Terms 50.3 provides that AWS may use and store the content processed by the listed AI Services — a list that names Amazon Polly with no tier qualifier, in the same sentence where AWS does qualify other products by tier. Opting out is done through an AWS Organizations opt-out policy, which means an organization administrator must act; there is no toggle inside Polly.
The Polly product page states the opposite in plain language: “Amazon Polly does not retain the content of your text submissions” (aws.amazon.com/polly, checked September 1, 2026).
And the FAQ frames it as consent-based — “You always retain ownership of your content and we will only use your content with your consent” — while Service Terms 50.3 obtains that consent by drafting it as your standing instruction. The same FAQ page also describes the opt-out mechanism two different ways in two different answers.
We are not going to reconcile a product page, an FAQ and a contract that disagree. The operative document in a dispute is the Service Terms, and the Service Terms describe default use of your content with an organization-level opt-out. If your compliance position depends on the product page’s sentence, get that in writing from AWS before you rely on it.
No retention period is stated anywhere we read. Whether generated audio, as distinct from submitted text, is stored or used for service improvement is also not stated: Service Terms 50.3 speaks of “AI Content”, the FAQ answer is scoped to “text inputs”, and neither settles the audio question.
Attribution and AI disclosure
No attribution or credit requirement appears on any official page we read — we searched the Customer Agreement, the FAQ, the pricing page, the Acceptable Use Policy and the Responsible AI Policy and found none. We record that as absent from the pages read rather than as an affirmative statement by AWS that none is required.
Likewise, no AI-disclosure or synthetic-audio labelling duty appears in those documents. The one watermarking obligation in Service Terms clause 50.12.6 is expressly scoped to other services, not Polly.
One clause worth reading before an IVR deployment
Service Terms 50.6 states that the AI Services “are not intended for use in, or in association with, the operation of any hazardous environments or critical systems that may lead to serious bodily injury or death”. AWS’s marketing FAQ, meanwhile, promotes Polly for public-announcement and notification use cases. Those two documents point different directions for anyone deploying to safety-adjacent announcements, and the contract is the one that governs.
Pros
- The lowest documented list price in our coverage — $4.00 per 1 million characters on Standard voices.
- Output ownership stated in a numbered contract clause rather than marketing copy (Service Terms 50.2).
- Caching and replay explicitly free, which materially changes IVR and prompt-library economics.
- A recurring monthly free allowance rather than a one-time trial, and it produces ordinary downloadable audio.
- The broadest documented regional footprint here on Standard voices, including GovCloud and a China region.
- No seats, no minimums, no annual commitment — billing scales down to nothing when you stop calling the API.
- Documented quotas and error codes throughout, which makes capacity planning possible without a sales conversation.
Cons
- Inputs are used to improve AWS models by default, with an opt-out that only an AWS Organizations administrator can set.
- AWS’s product page, FAQ and Service Terms give three different accounts of that data handling.
- No self-serve voice cloning at any price, and Brand Voice publishes no price, sample requirement or turnaround.
- Expressive SSML is documented as available only on the oldest, cheapest engine.
- The $100.00 Long-Form engine is documented in a single region.
- The generative engine produces no speech marks, ruling it out for lip-sync and word-highlighting.
- No published latency figure and no SLA for the service.
- Free-tier duration is stated inconsistently between the pricing page and the FAQ, and the newer six-month Free Plan is not reconciled with the twelve-month Polly allowances.
Best use cases
| Use case | Fit | Why |
|---|---|---|
| IVR and prompt libraries | Strong | Free replay of cached audio, lowest list price, broad regional availability |
| High-volume batch synthesis | Strong on paper | $4.00 per 1M characters and 10 tps default on Standard — but 1 tps on the premium engines |
| Accessibility and read-aloud features | Strong | Speech marks support word-level highlighting — on every engine except generative |
| Real-time voice agents | Mixed | Bidirectional streaming exists but is generative-only, 8 concurrent, with a 5-second idle timeout |
| Audiobooks and long narration | Weak on the premium path | Long-Form is one region and $100.00 per 1M characters, at 1 tps by default |
| Creator voiceover work | Weak | No editor, no cloning, no usable non-developer workflow |
And the inverse, because a recommendation without one is not worth much: if your project needs a voice that is recognisably yours, an editing timeline, or a documented consent process you can read before you commit, Polly does not serve it at any price on this rate card.
Alternatives
Organized by the reason you would leave.
- You need self-serve voice cloning. ElevenLabs documents instant cloning from its $6/month Starter plan and publishes its cloning rules openly (pricing, checked September 1, 2026).
- You want an editor rather than an API. Murf AI sells a timeline-based studio from $29/month, with a commercial grant you can quote in one sentence (pricing, checked September 1, 2026).
- You want comparable per-character economics without the AWS account. Google Cloud Text-to-Speech lists Standard and WaveNet at $0.000004 per character — $4 per million, matching Polly — with a 4-million-character monthly free allowance (pricing, checked September 1, 2026).
- You need the whole market in one view first. The ranking of ten providers compares pricing, free tiers, licensing and cloning side by side.
Direct comparison links
Head-to-head pages pairing Amazon Polly against individual competitors are on this site’s roadmap but not published yet, and we do not link to pages that do not exist. Until they are live, the flagship ranking carries the side-by-side pricing, licensing and feature tables for all ten providers we cover, including Polly.
Final verdict
Choose Amazon Polly when you are already operating inside AWS, your volume is large enough that a 25× price spread matters more than a user interface, and your legal review wants ownership language in a numbered clause. For cached, repeated audio — IVR trees, announcements, prompt libraries — the combination of the lowest documented list price and free replay is difficult to argue with on paper.
Do not choose it when you need a voice cloned without a sales conversation, an editing surface for non-developers, expressive control on a current-generation engine, or a contractual latency commitment. And do not choose it without first deciding, deliberately, whether your inputs may be used to improve AWS models — because the default answer is yes, and the opt-out lives at the organization level, not in Polly.
What would change this verdict is sound. Everything above is documentation. When our audio benchmark runs, this page gains measured comparisons across engines and against the rest of the market, and the assessment table gains scores that mean something. Until then this is a verdict about a rate card, a contract and a feature matrix — which is worth exactly as much as those three things, and no more.
Frequently asked questions
Did you actually test Amazon Polly’s voices?
No. We created no AWS account, made no API call and generated no audio for this review. Every statement here comes from AWS’s own documentation, read on September 1, 2026. Our listening benchmark is designed but not running (methodology).
Is Amazon Polly free?
There is a recurring monthly free allowance — 5 million characters on Standard voices, smaller amounts on the other three engines — and it produces ordinary downloadable audio rather than a locked preview. Its duration is the problem: the pricing page attaches “for the first 12 months” to three engines but not to Standard, while the FAQ describes the whole free tier as twelve months. Budget on the twelve-month reading.
Can I use Amazon Polly audio commercially?
AWS neither grants nor prohibits commercial use in those words; the word “commercial” does not appear on the pricing page or the FAQ. Your rights rest on Service Terms 50.2, which states that output you generate is Your Content. That is an ownership statement, and it is a numbered contract clause rather than marketing copy.
Does Amazon use my text to train its models?
By default, yes. Service Terms 50.3 provides for AWS using and storing content processed by its AI Services, and names Amazon Polly with no tier qualifier. The opt-out is an AWS Organizations policy set by an administrator. Note that AWS’s Polly product page states it does not retain your text submissions — the two documents disagree, and the Service Terms govern.
Can I clone my voice with Amazon Polly?
Not by yourself. None of Polly’s ten API operations creates a voice. Brand Voice is a custom engagement with the Polly team; AWS publishes no price, no sample-length requirement and no turnaround for it.
Which engine should I use?
That is a cost decision before it is a quality one: the spread runs from $4.00 to $100.00 per 1 million characters. Note that the engine parameter is optional and defaults to standard, so omitting it does not give you the newest model. We cannot yet tell you which sounds best, because we have not listened.
How fast is Amazon Polly?
Unknown from official sources. AWS publishes no latency figure and no SLA for Polly on any page we read, so we have nothing to quote — not even a vendor claim.
Will my Polly voice sound the same in a year?
AWS does not promise that. Its documentation states that updates to training data and the model “could result in slight variations to the way the voices sound”. For serialized content recorded over time, plan for that.
Sources and what we could not verify
Every changing fact on this page was read from official AWS sources on September 1, 2026. Where two official AWS pages disagree, both readings are published above and neither is averaged or silently reconciled.
| Source | Used for |
|---|---|
| aws.amazon.com/polly/pricing | Engine rates, GovCloud rates, free-tier allowances, billing basis, caching and replay |
| Amazon Polly FAQs | Free-tier duration wording, ownership answer, replay of static prompts, opt-out descriptions |
| AWS Service Terms | Clauses 50.1, 50.2, 50.3, 50.6, 50.10 — AI Services definition, ownership, data use, high-risk exclusion, indemnity list |
| AWS Customer Agreement | Section 12 definition of Your Content; clauses 6.1, 6.2, 6.4 |
| AWS Responsible AI Policy | The consent prohibition covering voice and likeness |
| AWS Free Tier Terms | Free Plan expiry, evaluation-purposes language, credit programme conditions |
| SynthesizeSpeech API reference | Engine parameter, valid values, default engine, output formats and sample rates |
| Standard voices | Concatenative description, voice inventory, regional list, feature list |
| Neural voices | AWS's quality claim, engine description |
| Generative voices | Runaway-generation safeguard wording, speech-marks absence, voice-stability warning |
| Long-form voices | Single-region availability, voice naming, stability warning |
| Supported SSML tags | Per-engine availability of expressive tags |
| API operations | The complete ten-operation surface — used to establish the absence of a create-voice call |
| Amazon Polly features | Brand Voice description and exclusivity wording |
| Amazon Polly product page | The retention statement that conflicts with Service Terms 50.3 |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026, recorded rather than guessed:
- Whether the Standard free tier is perpetual or twelve months. The pricing page omits the duration qualifier that the other three engines carry; the FAQ applies twelve months to all. Both readings published; neither reconciled.
- How the six-month AWS Free Plan interacts with Polly’s twelve-month allowances. No official page reconciles the two schemes.
- Brand Voice pricing, minimum commitment, sample requirements and consent terms. None is published anywhere; AWS routes the question to an account manager.
- Valid sample rates for MP3 and OGG output. Four official AWS pages give four different sets. Reported, not reconciled.
- MP3 bitrate for any engine. AWS documents sample rate in hertz only. We publish no bitrate rather than infer one.
- Any latency figure or SLA. Not published on any page we read.
- Data retention period. No duration is stated; the FAQ says AWS may store text inputs without a time limit.
- Whether generated audio, as distinct from submitted text, is used for service improvement. Service Terms 50.3 speaks of “AI Content”; the FAQ answer covers “text inputs”. Neither settles it.
- Voice-talent consent and licensing for the stock voices. No official page discloses how those voices were sourced.
- Whether Polly output is covered by AWS’s IP indemnity. We verified Polly is absent from the clause 50.10 indemnified list, but AWS makes no affirmative statement either way.
- The gender of the Turkish voice Filiz, and the correct long-form English voice names. AWS’s own pages contradict each other on both.
- Whether the standard engine has any deprecation plan. We found no sunset date, legacy label or end-of-life notice on any page read, including the full document-history table. Recorded as absent rather than confirmed non-existent.
Change log
September 1, 2026 — First publication. All prices, terms, quotas and feature statements read from official AWS sources on this date.
Published September 1, 2026 · Benchmark audio, measured costs and scores will be added with a dated entry here when the audio benchmark runs (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.