Google Cloud Text-to-Speech
Published September 3, 2026 · This page indexes everything this site has verified about Google Cloud and links to where each finding is set out in full. It is not a review — the review is. No listening test has been run, so nothing here says how Google Cloud sounds.
Google Cloud Text-to-Speech is metered cloud infrastructure rather than a consumer product — four families of voices behind REST and gRPC APIs with published quotas, no plans, no editor and no self-serve cloning — and this site’s #2 pick of the ten providers it documents.
The finding that matters most: Google Cloud Text-to-Speech is the verified per-character price floor of this market — $0.004 per 1,000 characters on Standard and WaveNet with 4 million free characters a month on each, matched only by Amazon Polly — but that floor buys infrastructure and nothing else: billing must be enabled before the first free character, there is no editor and no self-serve cloning, the same account runs up to $160 per million on Studio voices, and the newest Gemini-TTS family is token-billed with no published characters-per-token conversion, so its per-character cost cannot be computed at all.
Everything below is drawn from pages on this site, each of which carries its own sources and check dates (how we verify).
On this page
What it costs
Pay-as-you-go only, with no plans, seats or subscriptions: Standard and WaveNet list at $0.004 per 1,000 characters ($4 per million) with 4 million free characters a month on each, while Neural2 and Polyglot are $16 per million, Chirp 3: HD $30, Instant Custom Voice $60 with no free usage, and Studio voices $160 — a fortyfold spread inside one account, and Studio is the highest per-character text-to-speech rate in this site’s comparison set (verified August 11, 2026 on the review; pricing comparison restates the $4.00 floor, the $16.00 Neural2/Polyglot rate, $60.00 Instant Custom Voice and the $160.00 Studio ceiling from rates read September 1, 2026). The Gemini-TTS family bills in tokens instead — $0.50–$1.00 per million input text tokens and $10–$20 per million output audio tokens, with no free usage — and Google’s published conversion of 25 audio tokens per second normalizes output to roughly $0.015 per minute on the Flash models and $0.030 on Pro and 3.1 Flash, while per-character cost cannot be computed at all because no characters-per-text-token figure is published. Billing must be enabled before the first free character and overage bills automatically; the $300/90-day new-customer trial requires a payment method and carries no SLA or indemnity.
Every rate above is set against every other provider’s, normalized to one unit, in our pricing comparison.
Where Google Cloud’s own material disagrees with itself
These are contradictions inside the vendor’s published documentation, not disagreements between us and the vendor. We report both readings rather than choosing one, and each links to the page where we set it out in full. We have published 11 in total; the 10 with the most bearing on a decision are below, and the rest sit in the pages linked underneath.
- Google’s marketing and pricing pages disagree about the free allowances. The product page’s free-usage blurb says "The first 1 million characters for WaveNet voices are free each month" and "For Standard (non-WaveNet) voices, the first 4 million characters are free", while the pricing page — checked the same day — gives Standard and WaveNet 0–4M free each and Neural2, Studio, Chirp 3 and Polyglot 0–1M each. The site calls the marketing copy stale on WaveNet and wrong about non-WaveNet tiers sharing one band, treats the pricing page as canonical, attributes both and averages nothing. Checked August 11, 2026. Where we published this.
- Google publishes two conflicting language counts on a single page — "75+ languages and variants" in the body against "40+ languages" in its own navigation embedded in that same page — while the site’s own count of the enumerated documentation table gives 61 codes, or 53 base languages, matching neither advertised figure. Checked August 11, 2026. Where we published this.
- In the site’s recount of every vendor’s language claims, Google is one of two providers publishing two incompatible figures on one page ("75+ languages and variants" and "40+ languages"), against the site’s count of 61 codes and 53 distinct base languages in its enumerated list — so neither advertised figure matches its own table. Read September 2, 2026. Where we published this.
- Google’s overall catalog claim is contradicted by its own other surfaces: the product page advertises "380+ voices across 75+ languages and variants" while other Google pages still say "220+ voices / 40+ languages", which the site records as an unmaintained blurb rather than averaging the two. The site adds that its own scan of the voices table found far more voice rows than 380 across locales, leaving the count unverified by hand. Checked August 11, 2026. Where we published this.
- Google’s documentation disagrees with itself on whether Chirp 3: HD supports the ALAW audio encoding — the dedicated Chirp 3 page lists it and the voices overview denies it, both read on the same day. Recorded as unresolved rather than reconciled. Checked August 11, 2026. Where we published this.
- Google’s documentation disagrees with itself on Chirp 3: HD’s SSML support: "SSML support" is marked Preview on the dedicated page while the voices overview still states that SSML isn’t supported. The site records this explicitly as a live contradiction. Checked August 11, 2026. Where we published this.
- Google’s Gemini-TTS marketing and documentation disagree about availability: marketing describes the models as available "in 75+ locales" with no generally-available qualifier anywhere near it, while the documentation table’s launch-readiness column marks 24 of its 87 listed languages GA and 63 Preview — so roughly three quarters of the marketed coverage is preview. Read September 2, 2026. Where we published this.
- Google’s own terms leave the legal classification of this service unresolved: the Services Summary files Cloud Text-to-Speech under "AI/ML Services → Pre-Trained APIs", outside the named Generative AI Services list, while carrying a catch-all extending those terms to "any Generally Available generative AI features of a Service" without saying whether Chirp 3: HD or Gemini-TTS qualify. Section 20’s protections (explicit output-ownership language) and restrictions (under-18 audience, healthcare, prohibited use) may or may not attach, and no Google page resolves it. Checked August 11, 2026. Where we published this.
- The same unresolved classification is set out against Amazon Polly’s position: Cloud Text-to-Speech appears in Google’s Services Summary (last modified August 27, 2026) under Pre-Trained APIs rather than Generative AI Services, the section governing ownership and prohibited use applies to the latter, and the catch-all is not resolved on any Google page the site read. The general grant in Cloud Terms §1.1 permits use of the Services, and ownership is settled by §5.1 read with the definition of Customer Data, which reaches data derived through use of the Services — a position this site corrected on September 2, 2026 after initially recording it as unresolved. Checked September 1, 2026. Where we published this.
- Google’s authentication documentation and its own live discovery document disagree: the authentication page presents OAuth bearer tokens as the general REST pattern and shows no API-key header anywhere, while the live v1 discovery document declares an API key as a query parameter named key, described as required unless an OAuth token is supplied. Read September 2, 2026. Where we published this.
Dated and time-sensitive
Facts with a clock on them — model sunsets, policy revisions, market status. These are the entries most likely to be out of date first, which is why they carry their dates.
- 2026-06-01. The Google Cloud Platform Terms of Service carry a last-modified date of June 1, 2026. §5.1 is the ownership clause: "As between the parties, Customer retains all Intellectual Property Rights in Customer Data and Customer Applications." Source on this site.
- 2026-07-29. Google’s Service Specific Terms were last modified July 29, 2026 — the document carrying §17.a’s ban on using the service or its output "to develop a similar or competing product or service", §17.b’s model restrictions, §18’s commitment not to train on Customer Data without permission, and §20’s Generated Output, under-18 and healthcare provisions. The quotas page carries the same July 29, 2026 last-updated date. Source on this site.
- 2026-07-20. Checked against the July 20, 2026 revision of Google’s Generative AI Indemnified Services list: Text-to-Speech is not on it, so Google’s generative-output IP indemnity does not extend to this service. Separately, SLAs and Google’s indemnity do not apply during the $300 free trial. Source on this site.
- 2026-08-27. Google’s Services Summary — the document that classifies Cloud Text-to-Speech under Pre-Trained APIs rather than Generative AI Services — carries a last-modified date of August 27, 2026 and still carries that classification. Source on this site.
- 2026-09-01. Google’s documentation host moved: older citations pointing at cloud.google.com/text-to-speech/docs/… now redirect to docs.cloud.google.com. The pricing page and terms pages are unaffected. Source on this site.
- 2026-09-02. Correction published by this site: it had reported that whether you own audio generated with Google Cloud Text-to-Speech was unresolved, and that was wrong. Cloud Terms §5.1 gives the customer all Intellectual Property Rights in Customer Data, and Customer Data is defined to include "data that Customer or End Users derive from that data through their use of the Services" — which is what generated audio is. The site states it had located §5.1 but not followed through to the definition. Source on this site.
- 2026-09-02. The commercial-use guide carries the corrected position: Google publishes no clause permitting or prohibiting commercial use of synthesized audio, and the general grant in §1.1 permits use of the Services — but ownership is settled by §5.1 read with the definition of Customer Data. Silence on commercial use is recorded as neither permission nor prohibition. Source on this site.
- 2026-09-02. Google sells no dubbing product. The site enumerated Google’s sitemap — 180 sub-sitemaps — and searched every URL for "dub": exactly one matched, it sits under a /docs/private/ path, and it returns HTTP 404 on both of Google’s documentation hosts. Source on this site.
Every page on this site about Google Cloud
- Full review. Google Cloud Text-to-Speech Review 2026: Pricing, Licensing and Limits, Docs-Verified — the complete documentation-based assessment, with sources and dates against every claim.
- Alternatives. Google Cloud Text-to-Speech alternatives — organized by the reason you would leave, not as a ranked list.
- Head to head. Amazon Polly vs Google Cloud, Azure Speech vs Google Cloud, ElevenLabs vs Google Cloud, Google Cloud vs Murf AI.
- In our buying guides. Best AI Voice Generators in 2026: 10 Providers Compared, Best AI Dubbing Software 2026: Only Four Vendors Actually Sell It, Best Free AI Voice Generators 2026: What “Free” Actually Buys You, Most Realistic AI Voices: What Every Vendor Charges for the Claim, Best Multilingual AI Voice Generators: Every Count, Recounted, Real-Time AI Voice: Why the Latency Numbers Cannot Be Compared, Best AI Voice Cloning Software 2026: Price, Access and Consent Compared.
- In our comparisons of published data. AI Voice Generator Pricing Comparison: Every Rate Normalized, Best Text-to-Speech APIs: Limits, Transports and Auth Compared.
- In our explainers. Can You Use AI Voices Commercially? What Ten Vendors Actually Say, How AI Voice Cloning Works: What Ten Vendors Require of You, How AI Voice Generators Work, According to the Vendors Themselves.
- By what you are making. AI Voice for Audiobooks: Will Chapter 20 Match Chapter 1?, AI Voice for Business: What Survives a Procurement Review, AI Voice for eLearning: Courses Change, and the Audio Must Follow, AI Voice for Podcasts: Only Three Providers Do Two Speakers, AI Voice for YouTube: What You Must Disclose Before You Publish.
What we could not verify
Honesty about gaps beats a page that looks complete. Standing limits on everything above, as of September 3, 2026:
- How Google Cloud sounds. Our listening benchmark is designed but not running. Nothing on this page or on any page it links to ranks Google Cloud on audio quality, naturalness or speed.
- Anything Google Cloud does not publish. Every page indexed here is built from the vendor’s own documentation. Where that documentation is silent, we record the silence rather than estimating around it.
- Whether a documented feature works as documented. We read documentation; we did not create an account or make an API call.
- Currency. Each linked page carries its own check date. This index inherits them and does not re-verify on load.
Change log
September 3, 2026 — First publication. Assembled from pages already published on this site; no new vendor claim was introduced by this page.
Published September 3, 2026 · This page shows no audio and reports no listening results, because our benchmark has not run. Corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.