Azure Speech in Foundry Tools Review 2026
Published September 1, 2026 · Every fact below was read from Microsoft’s own pricing page, Learn documentation, responsible-AI pages and first-party pricing API on September 1, 2026, with the source linked beside the claim. Where a figure could not be read from the pricing page, this review says where it came from instead.
Quick verdict
Microsoft sells the largest voice catalog in this comparison behind two doors a self-serve buyer cannot open: a pricing page that displays no prices, and a cloning feature closed to anyone without a Microsoft account team. The access rule is stated plainly: “As a Limited Access feature, access to custom neural voice requires registration. Only customers managed by Microsoft, meaning those who are working directly with Microsoft account teams, are eligible for access” (limited access, checked September 1, 2026). That gate covers professional voice fine-tuning, personal voice and custom avatar alike.
The pricing page publishes no numbers at all. Every one of its 142 price cells renders as the literal string “$-”, and the page carries its own caveat: “Prices are estimates only and are not intended as actual price quotes.” The figures in this review come from Microsoft’s first-party Azure Retail Prices API, and the headline rate is independently corroborated by a Microsoft Learn document — see pricing.
What you get for that friction is depth. The catalog, language coverage, deployment options and enterprise plumbing are the broadest here. This is a platform bought through a procurement process, not a signup form.
Verdict basis: official documentation only, checked September 1, 2026 — no account, no API call, no listening test, no numeric score (how we verify).
| Best for | Enterprises already on Azure that need breadth, regional choice and container or on-premise deployment |
|---|---|
| Not best for | Any buyer without a Microsoft account team who needs voice cloning, or anyone who wants to read a price before committing |
| Price | $15.00 per 1M characters for standard neural; $22.00 per 1M characters for Neural HD. Neither figure appears on the pricing page — see sourcing note below |
| Free tier | 0.5 million characters per month on the F0 tier, Neural voices. Unused custom models are decommissioned after 7 days |
| Voice cloning | Closed by default. Requires Microsoft-managed customer status — not purchasable self-serve at any price |
| Product name | Currently “Azure Speech in Foundry Tools”. Formerly Azure AI Speech, and before that Azure Cognitive Services Speech |
| Test status | Documentation-verified. Our listening benchmark has not run, so there is no independent voice-quality score and no measured latency on this page. |
Every price, licensing term and feature claim below was read from Microsoft’s official pages on September 1, 2026, with the source linked beside the fact. Our standardized audio benchmark is designed but not yet running, so this review contains no listening-test claims, no naturalness verdicts and no scores.
Strongest advantages:
- The broadest documented catalog and language coverage in this comparison, across several voice tiers.
- A genuinely recurring free tier: 0.5 million characters per month, not a one-off trial.
- Container and on-premise deployment options that almost nothing else here offers.
- Enterprise plumbing throughout: regional choice, SDKs across major languages, batch synthesis.
- A published, dated responsible-AI position with an explicit access-control regime for cloning.
Important disadvantages:
- The pricing page renders no prices at all — 142 cells display a placeholder.
- Voice cloning is unavailable to any customer without a Microsoft account team.
- The product has been renamed twice and legacy naming is still live across SDKs, URLs and resource types.
- Voice and language counts contradict each other across pages, and sometimes within a single page.
- A much larger free allowance is still visible on the page inside a deprecated block, which is easy to misread as current.
- Unused custom models on the free tier are decommissioned after seven days.
On this page
What this product is called now
The current name is Azure Speech in Foundry Tools. We verified it on five separate Microsoft surfaces on September 1, 2026: the pricing page title and heading, the Learn overview, the SDK page, the quotas page and the AI code of conduct.
Microsoft documents the lineage in one place, verbatim: “Microsoft’s AI Platform has evolved from Azure AI Studio → Azure AI Foundry → to Microsoft Foundry (current). Similarly, our AI services portfolio evolved with the platform from Azure Cognitive Services → Azure AI Services → to Foundry Tools (current). Despite the platform evolution, the Azure resource type remains Microsoft.CognitiveServices/accounts.” So the earlier names you will find in older write-ups — Azure AI Speech, and before that Azure Cognitive Services Speech — refer to this product.
We could not establish a rename date, and we will not invent one. The documentation carrying the new name is stamped January 30, 2026, but that is a page metadata date, not a rename announcement. We searched the release-notes page in full and found no dated rebrand entry at all. Anyone citing a specific rename date should say where they got it, because Microsoft does not appear to publish one.
Two practical consequences. First, legacy naming is still live everywhere — SDK namespaces still read Microsoft.CognitiveServices.Speech, the container registry path still says azure-cognitive-services, the custom-voice documentation still sits at a /custom-neural-voice URL, the Azure resource type is unchanged, and the RBAC roles are still named “Cognitive Services Speech Contributor” and “Cognitive Services Speech User”. Your code will not stop working, but your search results will be confusing.
Second, there is a sub-rename worth knowing: “custom neural voice” is now “custom voice”. The documentation page is titled “What is custom voice?” and the phrase “neural voice” does not appear on it once — while its URL still contains custom-neural-voice. Prebuilt voices are called “Standard voice” in the documentation and “Neural” on the pricing page.
Who should use it — and who should look elsewhere
Choose Azure Speech if:
- You are already an Azure customer with an account team. That relationship is what unlocks cloning, and it is also how you will get a real price.
- You need breadth of voices and languages. Microsoft’s catalog is the largest documented here across several tiers, even though the exact counts disagree between pages.
- You need container or on-premise deployment. Microsoft documents container images for speech synthesis, which very few providers in this comparison offer at any price.
- You want a recurring free allowance to develop against. The F0 tier gives 0.5 million characters a month on Neural voices, and it recurs.
Look elsewhere if:
- You need to clone a voice and you are not a Microsoft-managed customer. There is no self-serve path at any price. ElevenLabs documents cloning from $6/month and Cartesia from $5/month (ElevenLabs pricing, Cartesia pricing, both checked September 1, 2026).
- You want to read a price before you commit. Microsoft’s own pricing page does not show one. Amazon Polly publishes $4.00 per million characters in plain text (pricing, checked September 1, 2026).
- You need per-character costs at the bottom of the market. At $15 per million characters, standard neural synthesis here lists at roughly four times Google’s and Amazon’s standard rates.
- You need voice and language counts you can quote. Microsoft’s own pages disagree — see voices and languages.
Current price, free plan, commercial rights
The short version, and an unusual sourcing note. Microsoft’s pricing page publishes labels, billing units, footnotes and free-tier allowances as ordinary text — but every actual price renders as “$-”. We counted 142 such cells, and the only dollar-and-digits string in the entire 646KB document is “$200”, which is free-account credit boilerplate. The page also carries its own caveat: “Prices are estimates only and are not intended as actual price quotes.”
The figures below therefore come from Microsoft’s own first-party Azure Retail Prices API, queried for the Foundry Tools service in the East US region on September 1, 2026, and reported exactly as returned. The headline rate is separately corroborated by a Microsoft Learn cost-estimation document that states in plain text: “Multiply the result by the unit price of $15 per million characters to estimate the monthly cost.” Two independent Microsoft sources agreeing is the reason we are comfortable publishing it at all.
| Meter | Price | Unit | Note |
|---|---|---|---|
| S1 Neural Text To Speech Characters | $15.00 | per 1M characters | Corroborated in a Learn document |
| Neural HD Text to Speech Characters | $22.00 | per 1M characters | Effective from March 1, 2026 |
| Free Text To Speech Characters | $0.00 | per 1M characters | The F0 allowance |
Microsoft’s own footnote sets out how the billing works: “Text to Speech: speech synthesis usage is billed per character. Avatar is billed per second. Training and model hosting is billed per second.” A separate footnote explains the tiering: “Selected text to speech voices are available via two model variants: Neural and NeuralHD.”
The free tier, and the larger number that is not it
The current free allowance is stated as literal text on the pricing page: “Text to Speech / (per character billing) / Neural / 0.5 million characters free per month.” It recurs monthly, and it is the only current text-to-speech free allowance.
There is a second, much larger figure on the same page, and it is not current. Inside a block Microsoft labels as deprecated — “These features are being deprecated and only available for existing customers to use” — a legacy free tier offers 5 million characters per month. Ten times the current allowance, on the same page, a scroll away. If you have seen Azure quoted as offering five million free characters, that is where the number comes from, and it is not available to a new customer.
Two F0 limits matter for planning, both quoted from the page’s own footnotes: “Unused models will be automatically decommissioned after 7 days”, and free speech-to-text hours are “shared between Standard and Custom” with batch unsupported.
What this review covers
This review covers Microsoft’s text-to-speech offering within Azure Speech in Foundry Tools: the current and legacy product naming, the published rates and where they actually come from, the free tier and its limits, the voice tiers, the access regime governing cloning, and the deployment and regional options. It touches speech-to-text only where the free tier is shared.
It does not evaluate how any Microsoft voice sounds. We created no Azure account, made no API call and generated no audio. Rows a review of ours normally carries once our benchmark runs — Tested, Model tested, Voice tested, Benchmark cost — are absent rather than filled with invented values.
One URL note, because it affects anyone checking our sources: the older pricing URL under /cognitive-services/speech-services/ now redirects to /pricing/details/speech/, and the delivered page’s own canonical link points there. We cite the destination.
Benchmark audio status
Samples are pending. Our standardized audio benchmark is designed but not running, so this page publishes no audio and no empty player (methodology). What we verified instead is documentation: the product naming across five surfaces, rates from two independent Microsoft sources, free-tier allowances and their footnotes, the limited-access regime, regional availability, and the voice-count discrepancies catalogued below, all read on September 1, 2026.
Assessment by dimension
No numeric scores appear below. The judgments are qualitative and drawn from what Microsoft documents.
| Dimension | What the documentation supports |
|---|---|
| Catalog breadth | Strong. The largest documented voice and language coverage here, across several tiers. |
| Pricing transparency | Weak. The pricing page renders no prices; usable figures require the retail API or a Learn document. |
| Cloning access | Closed. Limited-access registration restricted to Microsoft-managed customers; no self-serve path exists. |
| Free tier | Adequate and recurring at 0.5M characters/month, but easy to confuse with a deprecated 5M figure on the same page. |
| Deployment options | Strong. Containers and on-premise documented, with wide regional choice. |
| Documentation consistency | Weak. Voice counts, language counts, default parameters and region lists contradict each other, sometimes within one page. |
| Naming clarity | Weak. Two renames deep, with legacy identifiers still live throughout the platform. |
| Voice quality | Not assessed. We have run no listening test. |
| Latency | Not assessed. No figure we would publish. |
Voice cloning and the access gate
This is the single most consequential fact about Azure Speech for anyone shopping on capability, and it is stated without ambiguity in Microsoft’s responsible-AI documentation:
“As a Limited Access feature, access to custom neural voice requires registration. Only customers managed by Microsoft, meaning those who are working directly with Microsoft account teams, are eligible for access.”
That gate is not limited to the flagship cloning product. It covers professional voice fine-tuning, personal voice — including the demo in Speech Studio — and custom avatar.
What this means for you: if you do not have a Microsoft account team, you cannot clone a voice on this platform. Not on a higher tier, not by paying more, not by agreeing to extra terms. It is an eligibility question rather than a pricing question, and no amount of budget resolves it. For an independent developer or a small studio, that single paragraph removes Azure from consideration for any cloning work.
It is worth saying that this is a defensible policy rather than an oversight — gating synthetic voice creation behind a managed relationship is a real abuse control, and Microsoft publishes its reasoning. But a buyer needs to know the door is locked before they spend a week evaluating the room.
Voice quality and naturalness
We will not tell you how Microsoft’s voices sound. No listening test has been run for this review.
Microsoft markets several quality tiers, including high-definition voices with expressive controls. Those are vendor positions and we do not restate them as findings. What documentation settles is the tier structure, the price difference between Neural and Neural HD, and the regional availability of each — and on the last of those, Microsoft’s own pages disagree, as described next.
Voices and languages: the counts do not agree
Microsoft’s catalog is the broadest here. How broad is genuinely unclear from official sources, and the disagreements are not minor.
- Total preset voices and languages: three official pages give three different figures — “over 400 options in more than 140 languages and locales” on one, “more than 500” voices on another, and a third figure elsewhere.
- High-definition voice count: a single page states three different totals in three places.
- Language counts for individual voice families: one page gives two different numbers for the same family in its feature table and its notes; another does the same for personal voice.
- A default parameter value is stated one way in prose and another in the parameter table, on the same page.
- Regional availability of HD voices differs between the regions page and the feature documentation.
We are not going to pick the most flattering number, or average them. If a specific locale or a specific region decides your purchase, verify it against the region and language pages for the exact voice tier you intend to buy, and expect to confirm it with Microsoft.
API, SDKs, deployment
This is the strongest part of the offering and the reason enterprises choose it. Microsoft documents Speech SDKs across the major languages, batch synthesis for long-form work, wide regional deployment, and container images that allow synthesis to run inside your own infrastructure — an option almost nothing else in this comparison provides.
Note that the SDK namespaces and container registry paths still carry the pre-rename Cognitive Services naming, as described in naming. That is cosmetic, but it means package searches and older code samples remain valid.
Generation speed and latency
We publish no latency figure for Azure Speech. We measured none, and we did not find a Microsoft figure in a form we would cite as a specific documented claim for text-to-speech.
This section will carry measured figures when our benchmark runs.
Pricing and normalized cost
What you pay, and what you get
Billing is per character, so the rate is the effective cost above the free allowance: cost = (characters ÷ 1,000,000) × the tier rate. A 100,000-character chapter costs $1.50 on standard neural and $2.20 on Neural HD, at the rates returned by Microsoft’s retail API on September 1, 2026. Those are computations from published rates, not measurements of a workload we ran.
The important limitation: you cannot check the price on the price page
Every figure above required going around Microsoft’s pricing page rather than reading it. For a buyer that is more than an inconvenience: it means the number you take to a budget meeting cannot be sourced to the page your finance team will check. We would put the Learn cost-estimation document in front of them instead, since it states the $15 rate in plain prose, and treat the retail API as the authority for anything the Learn document does not cover.
Who gets the best value
- Enterprises with an Azure commitment: strong, because the negotiated rate and the account relationship are the actual product here.
- Developers on the free tier: reasonable. Half a million characters a month recurring is enough to build against.
- High-volume price-sensitive buyers: weak. At $15 per million characters, standard synthesis lists well above the cloud incumbents.
Commercial usage, licensing, privacy
Azure Speech is governed by Microsoft’s general product and privacy terms rather than a speech-specific contract, with the service appearing under the Microsoft Azure Core Services heading in the licensing terms. Microsoft’s licensing-terms pages are served in a way that blocks a plain HTTP client and requires a rendering fetch, which we note so you know why we cite them sparingly here.
The layer that is speech-specific, well documented and genuinely important is the responsible-AI regime: the limited-access rule quoted in cloning, and the AI code of conduct that Microsoft states “unifies and replaces the previous codes for Microsoft Generative AI Services, Azure Face in Foundry Tools, and Azure Speech in Foundry Tools text to speech.” That consolidation is itself the clearest dated evidence we found that the older product name was in use as recently as 2025.
Because the general terms cover many services at once, a buyer evaluating output ownership or data handling for speech specifically should confirm the position with Microsoft against the current product terms rather than rely on a summary — including ours. We would rather point you at that step than paraphrase a multi-service contract into a sentence it does not support.
Pros
- The broadest documented voice and language catalog in this comparison.
- A recurring free tier of 0.5 million characters per month rather than a one-off trial.
- Container and on-premise deployment, rare among the providers here.
- Wide regional choice and mature enterprise SDK coverage.
- Batch synthesis for long-form work.
- A published, dated responsible-AI position with a real access-control regime for cloning.
- The headline rate is corroborated by two independent Microsoft sources.
Cons
- The pricing page renders no prices at all — 142 placeholder cells.
- Voice cloning is closed to any customer without a Microsoft account team, at any price.
- Standard neural lists at roughly four times the standard rates of Google and Amazon.
- Two product renames deep, with legacy naming still live across SDKs, URLs, RBAC roles and resource types.
- Voice, language and region counts contradict each other across pages and within single pages.
- A deprecated 5-million-character free tier still displays on the pricing page and is easy to mistake for the current one.
- Unused custom models on the free tier are decommissioned after seven days.
- The licensing pages are not readable by ordinary tooling, which makes independent verification harder.
Best use cases
| Use case | Fit | Why |
|---|---|---|
| Enterprise deployment on Azure | Strong | Account relationship, regional choice, containers, mature SDKs |
| Multilingual coverage at breadth | Strong on paper | Largest catalog documented here, though the counts disagree |
| On-premise or air-gapped synthesis | Strong | Container images are documented, which is rare in this market |
| Any cloning project without an account team | Blocked | Limited-access eligibility, not a pricing tier |
| Independent developers and small studios | Weak | No self-serve cloning, no readable prices, rate well above the incumbents |
| Quick evaluation before purchase | Weak | You cannot read a price on the pricing page |
And the inverse, because a recommendation without one is not worth much: if your buying process is a card and an afternoon rather than a procurement cycle and an account manager, this platform is not built for you — regardless of how good the catalog is.
Alternatives
Organized by the reason you would leave.
- You need self-serve voice cloning. ElevenLabs from $6/month, Cartesia from $5/month (both checked September 1, 2026).
- You need a published, readable price. Amazon Polly publishes $4.00 per million characters in plain text (pricing, checked September 1, 2026).
- You need the lowest verified per-character rate with a large free tier. Google Cloud Text-to-Speech lists $4 per million characters with four million free monthly (pricing, checked September 1, 2026).
- You need on-premise without a vendor relationship. Resemble AI publishes MIT-licensed models you can self-host.
- You want the whole market in one view. The ranking of ten providers compares pricing, free tiers, licensing and cloning side by side.
Direct comparison links
Head-to-head pages pairing Azure Speech against individual competitors are on this site’s roadmap but not published yet, and we do not link to pages that do not exist. Until they are live, the flagship ranking carries side-by-side pricing, licensing and feature tables for all ten providers we cover.
Final verdict
Choose Azure Speech in Foundry Tools when you are buying as an enterprise, on Azure, with an account team who can quote you a real number and unlock the features that matter. On those terms it is a serious platform: the broadest catalog here, container and on-premise deployment, regional depth, and a responsible-AI regime that is published rather than implied.
Do not choose it when your buying process does not include a Microsoft account manager. Cloning is an eligibility gate rather than a price tier, the pricing page will not show you a number, the standard rate lists well above the cloud incumbents, and the documentation contradicts itself often enough that you will end up confirming basic facts by email. None of that is fatal to a large enterprise. All of it is fatal to a quick evaluation.
What would change this verdict is a pricing page that renders prices — and, separately, sound. Fixing 142 placeholder cells would move this offering up the comparison without changing a single feature. When our audio benchmark runs, this page gains measured comparisons and the assessment table gains scores that mean something. Until then this is a verdict about documentation, an access policy, and a price list we had to reach through an API to read.
Frequently asked questions
Did you actually test Azure Speech’s voices?
No. We created no Azure account, made no API call and generated no audio for this review. Every statement here comes from Microsoft’s own documentation and first-party pricing API, read on September 1, 2026. Our listening benchmark is designed but not running (methodology).
What is this product called now?
Azure Speech in Foundry Tools, verified on five Microsoft surfaces. It was previously Azure AI Speech, and before that Azure Cognitive Services Speech. Microsoft does not appear to publish a rename date, and we do not invent one.
How much does Azure text-to-speech cost?
$15.00 per million characters for standard neural and $22.00 per million for Neural HD, from Microsoft’s retail pricing API. The $15 figure is separately confirmed in a Microsoft Learn document. Neither number appears on Microsoft’s own pricing page, where all 142 price cells render as a placeholder.
Why does the pricing page show no prices?
We do not know, and we will not speculate. What we can report is what we observed: 142 cells rendering the literal string “$-”, prices injected client-side, and the page’s own caveat that prices are estimates rather than quotes.
Can I clone a voice on Azure?
Only if you are a Microsoft-managed customer. The limited-access rule restricts registration to “customers managed by Microsoft, meaning those who are working directly with Microsoft account teams”, and it covers professional voice fine-tuning, personal voice and custom avatar. There is no self-serve route at any price.
Is there a free tier?
Yes: 0.5 million characters per month on Neural voices, recurring. Be careful of a 5-million-character figure elsewhere on the same page — it sits inside a block Microsoft labels as deprecated and available only to existing customers.
How many voices and languages are there?
Microsoft’s pages disagree. Different official sources give “over 400 options in more than 140 languages and locales” and “more than 500” voices, and individual voice families are given two different language counts on a single page. Verify the specific locale you need.
Will my old Azure Speech code still work?
The naming changed but the plumbing did not: SDK namespaces, the container registry path, the RBAC role names and the Azure resource type all still use the pre-rename Cognitive Services identifiers.
Sources and what we could not verify
Every changing fact on this page was read from official Microsoft sources on September 1, 2026. Because the pricing page renders no numbers, rates come from Microsoft’s own retail pricing API, with the headline rate corroborated by a Microsoft Learn document. Where Microsoft’s pages contradict each other, all readings are reported and none is reconciled.
| Source | Used for |
|---|---|
| Azure Speech pricing page | Product name, billing-unit footnotes, free-tier allowances, the deprecated legacy tier, the estimates caveat — and the 142 placeholder price cells |
| Azure Retail Prices API (Foundry Tools, East US) | Standard neural, Neural HD and free-tier meter rates |
| Speech service overview | Current product name and service description |
| Limited access for text to speech | The cloning access gate and the features it covers |
| Navigating from classic | The documented platform and portfolio naming lineage |
| Custom voice overview | The custom-neural-voice to custom-voice sub-rename |
| Speech service release notes | Searched in full for a dated rebrand entry; none found |
| Quotas and limits | Current product naming and service limits |
| High-definition voices | Voice tiers and the internally inconsistent counts |
| Regions | Regional availability, including the HD region discrepancy |
| AI code of conduct | The consolidation statement naming the former product |
What we could not verify
Honesty about gaps beats a page that looks complete. Genuinely unresolved as of September 1, 2026, recorded rather than guessed:
- Any price on Microsoft’s own pricing page. All 142 price cells render as a placeholder. Our figures come from Microsoft’s retail API and a Learn document instead.
- The date of the rename. No dated rebrand entry exists in the release notes; the January 30, 2026 stamp is page metadata, not an announcement.
- The total number of voices and languages. Three official pages give three different figures.
- Language counts for individual voice families. Two families are given conflicting counts within a single page each.
- Regional availability of HD voices. The regions page and the feature page disagree.
- A default parameter value stated one way in prose and another in the table on the same page.
- The full licensing position for speech specifically. Microsoft’s licensing-terms pages are not readable by ordinary tooling, and the terms cover many services at once; we point buyers at that step rather than paraphrase it.
- Any vendor latency figure for text-to-speech in a form we would publish.
- Custom-voice pricing. Not reachable without the access registration described above.
Change log
September 1, 2026 — First publication. Product naming verified across five Microsoft surfaces; rates taken from Microsoft’s first-party retail pricing API and corroborated against a Microsoft Learn document, because the pricing page renders no numbers.
Published September 1, 2026 · If the pricing page begins rendering prices, the figures here will be re-sourced to it and re-dated. Benchmark audio, measured costs and scores will be added with a dated entry when the audio benchmark runs (methodology). Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.