Azure Speech in Foundry Tools vs Google Cloud Text-to-Speech 2026
Published September 2, 2026 · This is a documentation-based comparison. Every figure below was read from the two vendors’ own sources on September 1, 2026 and carries that check date. No listening test has been run, so nothing here compares how the two providers sound.
The short answer
Google is cheaper on every axis a buyer can check, and it is the only one of the two whose price you can read on its own pricing page. Google’s Standard and WaveNet voices list at $4 per million characters with four million free every month. Azure’s standard neural rate is $15 per million with half a million free — and that $15 came from Microsoft’s retail pricing API and a Learn document, because all 142 price cells on the Azure Speech pricing page render as the placeholder “$-” (Google pricing, Azure pricing page, both checked September 1, 2026).
Azure’s case is breadth and deployment. The largest documented voice catalog in our comparison set, plus container and on-premise deployment that Google does not offer for this service. If audio must never leave your infrastructure, that decides it regardless of price.
Neither sells self-serve cloning — Google’s is allow-listed with a published $60 per million rate, Azure’s is restricted to Microsoft-managed customers at any price.
Comparison basis: official documentation only, checked September 1, 2026 — no accounts created, no API calls made, no audio generated, no scores assigned (how we verify).
| Azure Speech in Foundry Tools | Google Cloud Text-to-Speech | |
|---|---|---|
| Standard rate | $15.00 per 1M characters — not published on the pricing page | $4 per 1M characters (Standard, WaveNet) |
| Premium rates | $22.00 per 1M (Neural HD) | $16 (Neural2, Polyglot); $30 (Chirp 3: HD); $160 (Studio) |
| Where the price comes from | Microsoft’s retail pricing API, corroborated by a Learn document | Stated in plain text on the pricing page |
| Free tier | 0.5M characters/month, Neural, recurring | 4M characters/month on Standard and WaveNet; 1M on premium families |
| Billing units | Per character throughout | Per character, except the Gemini-TTS family which bills per token |
| Voice cloning | Closed. Microsoft-managed customers only | Allow-listed via sales; $60 per 1M characters published |
| Deployment | Cloud, plus documented containers and on-premise | Cloud only |
| Catalog breadth | The broadest here, though the counts contradict across pages | Several documented voice families |
| Output ownership | General Microsoft product terms; no speech-specific clause found | Unresolved. Listed under Pre-Trained APIs, not Generative AI Services |
| Audio quality | Not compared. Our listening benchmark has not run, so this page assigns no naturalness verdict to either provider. | |
On this page
Winners by category
Every winner below is a documentation-based judgment from published prices, terms and feature documentation. None is a test result, and the categories that require listening have no winner.
| Category | Winner | Why |
|---|---|---|
| Price | Google Cloud | $4 against $15 per million characters on the standard tiers |
| Pricing transparency | Google Cloud | Rates are text on the page; Azure’s pricing page renders 142 placeholder cells |
| Free tier | Google Cloud | 4 million characters a month against 0.5 million — eight times larger |
| Deployment options | Azure | Containers and on-premise are documented; Google offers neither for this service |
| Catalog breadth | Azure | The largest documented coverage in our comparison set |
| Cloning access | Google Cloud | Both are gated, but Google publishes a price and does not restrict eligibility to managed customers |
| Billing-unit consistency | Azure | Per character throughout; Google adds a token-billed family that cannot be compared without measuring |
| Product-name stability | Google Cloud | Azure has been renamed twice, with legacy identifiers still live throughout the platform |
| Documentation consistency | Google Cloud | Azure’s own pages give conflicting voice counts, language counts and region lists |
| Naturalness | No winner | Requires listening. Our benchmark has not run |
| Pronunciation and difficult text | No winner | Requires listening. Our benchmark has not run |
Price and the pricing-page problem
Google is 3.75 times cheaper on the standard tiers: $4 against $15 per million characters. Even Azure’s cheapest rate sits close to Google’s mid tier — Neural2 and Polyglot at $16 per million — so a buyer choosing Azure on price alone has no case.
Above that the picture is less one-sided. Azure’s premium Neural HD is $22 per million, comfortably below Google’s Chirp 3: HD at $30 and far below Google’s Studio voices at $160 per million — the most expensive per-character rate in our entire comparison set. If your requirement is a premium voice tier, Azure is the cheaper premium option.
The structural difference is that only one of these two publishes its prices. Microsoft’s pricing page renders labels, billing-unit footnotes and free-tier allowances as text, and every actual price as the literal string “$-” — 142 cells of it. The page also states: “Prices are estimates only and are not intended as actual price quotes.” The $15 and $22 figures come from Microsoft’s own first-party Azure Retail Prices API, with the $15 rate independently corroborated by a Microsoft Learn document stating: “Multiply the result by the unit price of $15 per million characters to estimate the monthly cost.”
Google adds a different complication: its Gemini-TTS family is billed in tokens rather than characters — for instance $0.50 per million input text tokens and $10.00 per million output audio tokens on Gemini 2.5 Flash TTS, footnoted “Audio tokens correspond to 25 tokens per second of audio”. That family cannot be compared to a per-character rate without measuring real usage, and it sits outside Google’s free tier.
Free tiers
Google’s is eight times larger: four million characters a month on Standard and WaveNet, plus one million on the premium families, against Azure’s half a million on Neural. Both recur monthly and both produce ordinary usable audio.
Each carries one condition worth knowing. Google requires a billing account before the first free character: “You must enable billing to use Text-to-Speech, and will be automatically charged if your usage exceeds the number of free characters allowed per month.” Azure decommissions unused custom models after seven days, and displays a much larger legacy allowance — five million characters a month — inside a block it labels as deprecated and available only to existing customers. A new Azure customer cannot have that five million figure, however prominently it appears on the page.
Voice cloning
Neither is self-serve, and the gates differ in kind.
Azure’s is eligibility-based and absolute: “As a Limited Access feature, access to custom neural voice requires registration. Only customers managed by Microsoft, meaning those who are working directly with Microsoft account teams, are eligible for access.” No tier and no budget opens it if you lack that relationship, and the gate covers professional voice fine-tuning, personal voice and custom avatar alike.
Google’s is an allow-list with a published price: Instant Custom Voice costs $60 per million characters, and “Access to Instant Custom Voice is restricted to allow-listed users. To request access, contact a member of the sales team.” You can at least budget for it before asking.
If self-serve cloning is the requirement, neither of these is the answer. ElevenLabs documents it from $6 a month and Cartesia from $5 (both checked September 1, 2026).
Deployment, breadth and naming
Azure’s genuine advantage is deployment. Microsoft documents container images and on-premise deployment for speech synthesis. Google does not offer an equivalent for this service. For a buyer with a data-residency or air-gap requirement, that single capability outweighs a 3.75-fold price difference, because the cheaper option cannot do the job at all.
Azure also documents the broadest catalog in our comparison set — though how broad is genuinely unclear from official sources. Three Microsoft pages give three different totals for voices and languages, one page states three different high-definition voice counts, and regional availability of HD voices differs between the regions page and the feature documentation.
One naming note that costs research time rather than money. The product is currently Azure Speech in Foundry Tools, formerly Azure AI Speech and before that Azure Cognitive Services Speech. We could not establish a rename date — the documentation carrying the new name is stamped January 30, 2026, but that is page metadata and a full search of the release notes found no dated rebrand entry. Legacy naming remains live across SDK namespaces, the container registry path, RBAC roles and the Azure resource type, so older code and samples remain valid while search results stay confusing.
Ownership: two kinds of unresolved
Neither vendor gives you a speech-specific ownership clause, and they fail to in different ways.
Google’s is a classification problem. Cloud Text-to-Speech appears in Google’s Services Summary under “Pre-Trained APIs”, not under “Generative AI Services” — and the Service Specific Terms section governing ownership applies to the latter. Its own catch-all extends that heading to “any Generally Available generative AI features of a Service” without saying whether Chirp 3: HD or Gemini-TTS qualify. No Google page resolves it. The general grant in Cloud Terms §1.1 permits use of the Services; a specific ownership assignment for synthesized audio we did not find.
Azure’s is a breadth problem. The service is governed by Microsoft’s general product and privacy terms rather than a speech-specific contract, appearing under the Microsoft Azure Core Services heading. Those terms cover many services at once, and Microsoft’s licensing-terms pages are served in a way that blocks ordinary tooling, which makes independent verification harder than it should be.
What this means for you: if your contract with a client requires you to warrant that you own delivered audio, neither of these two hands you that warranty in a clause you can quote. Amazon Polly does, in Service Terms 50.2 — which is a real reason to consider it alongside these two.
Voice quality, pronunciation and naturalness
We have not compared how these two providers sound.
Our standardized listening benchmark is designed but not running. Until it does, this page carries no naturalness verdict, no pronunciation comparison and no scores for either provider. Both vendors publish quality claims of their own; those belong to them, and we do not repeat them as findings.
Choose Azure if · Choose Google if
Choose Azure Speech in Foundry Tools if:
- You need container or on-premise deployment. This is the decisive capability, and Google has no equivalent here.
- You need the broadest catalog, accepting that the exact counts are contested across Microsoft’s own pages.
- Your premium tier matters more than your standard tier. Neural HD at $22 per million undercuts Google’s Chirp 3: HD at $30 and its Studio voices at $160.
- You have a Microsoft account team, which is what unlocks cloning and a negotiated price.
Choose Google Cloud Text-to-Speech if:
- Price matters. $4 against $15 per million characters on the standard tiers.
- You want a free tier you can actually build on. Four million characters a month against half a million.
- You need to read a price on the vendor’s pricing page. Azure’s does not show one.
- You want cloning priced transparently even if gated — $60 per million published, against an eligibility gate no budget opens.
Sources and what we could not verify
Every figure on this page was read from the two vendors’ official sources on September 1, 2026. Full detail is in the individual reviews: Azure Speech in Foundry Tools review and Google Cloud Text-to-Speech review. How we verify anything is described on the methodology page.
| Source | Used for |
|---|---|
| Azure Speech pricing page | Billing-unit footnotes, free-tier allowances, the deprecated legacy tier, and the 142 placeholder cells |
| Azure Retail Prices API (Foundry Tools, East US) | The $15.00 and $22.00 per-million-character rates |
| Azure limited access | The cloning eligibility gate |
| Azure Speech overview | Current product naming and service description |
| Azure regions | Regional availability, including the HD discrepancy |
| Google Cloud TTS pricing | Per-character and per-token rates, free-tier bands, billing-account requirement |
| Google Cloud Platform Terms | The §1.1 services grant |
| Google Services Summary | Cloud TTS classified under Pre-Trained APIs (last modified August 27, 2026) |
| Instant Custom Voice docs | The sales allow-list restriction on cloning |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026:
- How either provider sounds. No listening test has been run.
- Any Azure price on Microsoft’s own pricing page. All 142 price cells render as a placeholder.
- Azure’s total voice and language counts. Three official pages give three different figures, and one page contradicts itself.
- The date of Azure’s rename. No dated rebrand entry exists in the release notes.
- Whether Google’s Generative AI ownership terms reach Cloud Text-to-Speech. It is classified under Pre-Trained APIs and the catch-all is unresolved on Google’s own pages.
- The full licensing position for Azure speech specifically. Microsoft’s licensing-terms pages are not readable by ordinary tooling and cover many services at once.
- Azure custom-voice pricing. Not reachable without the access registration.
- Any per-character comparison against Google’s Gemini-TTS models. They bill in tokens and no conversion is published.
- Any vendor latency figure for text-to-speech on either side, in a form we would publish.
Change log
September 2, 2026 — First publication as a documentation-based comparison. All figures read from official vendor sources on September 1, 2026 and dated accordingly.
Published September 2, 2026 · When our audio benchmark runs, this page gains a head-to-head listening test using the same prompts and documented conditions for both providers, and the naturalness and pronunciation rows gain real winners (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.