OpenAI
Published September 3, 2026 · This page indexes everything this site has verified about OpenAI and links to where each finding is set out in full. It is not a review — the review is. No listening test has been run, so nothing here says how OpenAI sounds.
OpenAI’s speech API: three dedicated text-to-speech models — gpt-4o-mini-tts, tts-1 and tts-1-hd — called on the same key, bill and SDK as the rest of the OpenAI platform, with delivery steered by a plain-language `instructions` parameter rather than SSML, and custom voices available only through a sales conversation (the review).
The finding that matters most: OpenAI pairs the best data and ownership paperwork in this site’s coverage — API content contractually excluded from training with the speech endpoint named, plus an express assignment of output rights — with a price a buyer cannot forecast: the current model bills audio output in tokens with no published conversion, the older models are billed per character in the same table, and there is no free tier to measure against, which the review calls "the single biggest obstacle to evaluating OpenAI’s speech offering on paper" (the review).
Everything below is drawn from pages on this site, each of which carries its own sources and check dates (how we verify).
On this page
What it costs
As read from OpenAI’s developer pricing page on September 1, 2026 and cross-checked against OpenAI’s own Markdown twin of that page the same day, which agrees exactly: gpt-4o-mini-tts is $12.00 per 1M tokens of audio output plus $0.60 per 1M tokens of text input; tts-1 is $15.00 per 1M characters; tts-1-hd is $30.00 per 1M characters. There is no free tier of any kind for speech — the only provider among the ten this site covers with no free path to a first generated file — and no per-minute figure is published for any speech model, the per-minute column on OpenAI’s own table being populated only for transcription (the review). pricing comparison (read September 1, 2026) normalizes tts-1 to $15.00 and tts-1-hd to $30.00 per million characters as published, and keeps gpt-4o-mini-tts out of the main table because no token-to-character factor is displayed for it anywhere on OpenAI’s pricing page; the site’s earlier flagship figures, checked August 11, 2026, were $0.015 and $0.030 per 1,000 characters (/best-ai-voice-generators). OpenAI’s Batch API is documented at "50% lower costs" with higher rate limits and a 24-hour turnaround (elearning page).
Every rate above is set against every other provider’s, normalized to one unit, in our pricing comparison.
Where OpenAI’s own material disagrees with itself
These are contradictions inside the vendor’s published documentation, not disagreements between us and the vendor. We report both readings rather than choosing one, and each links to the page where we set it out in full. We have published 8 in total, and all of them are below.
- OpenAI’s text-to-speech guide states that "Our usage policies require you to provide a clear disclosure to end users that the TTS voice they are hearing is AI-generated and not a human voice" — but the site went to the usage policies to read the requirement in its own words and did not find it there. The guide asserts an obligation and attributes it to a document that, as far as the site could establish, no longer contains it. Where we published this.
- Three speech models sit in one price table under two incompatible billing units. The section header states "Prices per 1M tokens unless noted", and the string "/ 1M characters" appears only as a per-cell override on the two tts-1 rows. Because OpenAI publishes no token-to-character or token-to-second conversion for audio output, its own three models cannot be compared against each other on the published figures. Where we published this.
- The tts-1 lineage is described three different ways across three official OpenAI pages, all current: tts-1 and tts-1-hd are absent from the "Speech generation" section of the models index, present with their own live model pages, and present with live price rows on the pricing page. The site reports the inconsistency rather than picking the tidiest reading. Where we published this.
- OpenAI’s guide sends parameters its own API reference does not document. The custom-voice example posts "language": "fr" and "format": "wav" to the speech endpoint, while the reference’s body-parameter list contains only input, model, voice, instructions, response_format, speed and stream_format — and a case-insensitive search of that reference for "language" returns nothing. Where we published this.
- The only enumerated language list on OpenAI’s text-to-speech guide — 57 language names — is presented as Whisper’s, and Whisper is speech recognition: it describes what OpenAI can transcribe, not what it can speak, and its authoritative version lives in a GitHub README rather than the API documentation. OpenAI publishes no text-to-speech language list of any kind, making it the only provider in the site’s coverage with none. Where we published this.
- Two different input caps are published for the same request. The speech endpoint reference states verbatim "The maximum length is 4096 characters", while the gpt-4o-mini-tts model page states a maximum of 2,000 input tokens. Where we published this.
- The custom-voices section of the guide closes with "Refer to the Text-to-Speech Supplemental Agreement for additional terms of use", but that sentence carries no hyperlink — the site checked the page markup — and the document is not among OpenAI’s published policy list. A buyer therefore cannot read the terms governing custom voices before contacting sales. Where we published this.
- OpenAI’s regional endpoints do not move the processing for text-to-speech. Ten regional hostnames are published, but for the audio endpoint group containing text-to-speech most regions offer regional storage only — only the United States and Europe show processing. Choosing a Tokyo endpoint for data residency does not put the synthesis compute in Tokyo. Where we published this.
Dated and time-sensitive
Facts with a clock on them — model sunsets, policy revisions, market status. These are the entries most likely to be out of date first, which is why they carry their dates.
- 2026-07-20. OpenAI’s dated deprecation notice records that on this date developers using legacy audio, realtime and transcription model families and snapshots were notified of their deprecation and removal from the API. A material share of the rows on the live audio price table therefore already carries an end date. Source on this site.
- 2026-08-11. OpenAI Services Agreement clauses 4.1, 4.2, 4.4, 9.1 and 9.2 were verified on this date (ownership assignment, training exclusion, non-uniqueness caveat, reservation of rights). Every contract clause quoted in the review’s licensing section carries this date rather than the review’s later check date. Source on this site.
- 2026-09-01. OpenAI’s policy pages returned an access error to every method tried, so the contract clauses could not be re-read; the marketing pricing page also returned an access error and could not be cross-checked. The developer pricing page was reachable and is the source for every published figure. The absence of a free tier is likewise established from the readable developer pricing page and the speech guide rather than the policy pages. Source on this site.
- 2023-03-01. The date OpenAI itself names in its no-training statement: "Your data is your data. As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)", with the speech endpoint listed by name in the per-endpoint data table. Source on this site.
- 2026-01-01. The site’s flagship comparison records the OpenAI Services Agreement clause 4.1 (express assignment of output rights) as effective January 1, 2026, checked August 11, 2026. Source on this site.
- 2026-06-12. OpenAI’s service terms carry an update date of June 12, 2026 per the flagship comparison (checked August 11, 2026), which records that ChatGPT’s free voice output is non-commercial and may not be repackaged as audio files under Service Terms §8. The later OpenAI review (the review) states this distinction as one to check rather than as a verified quotation, because the policy pages were unreachable on September 1, 2026. Source on this site.
- 2026-06-30. OpenAI publishes a SOC 2 report covering "July 1, 2025 to June 30, 2026", alongside ISO/IEC 27001:2022 "for OpenAI’s API, ChatGPT Enterprise, and ChatGPT Edu services". The site notes OpenAI is the only vendor in its coverage that dates its report period, and that the ISO scope names the API specifically (read September 2, 2026). Source on this site.
- 2026-09-01. First publication of the OpenAI review. Prices, model status, limits and guide statements read from official OpenAI developer sources on this date; contract clauses carry their August 11, 2026 verification date. No listening test, no measured latency and no numeric score appear on the page. Source on this site.
Every page on this site about OpenAI
- Full review. OpenAI Text-to-Speech Review 2026: Pricing, Cloning and Terms, Docs-Verified — the complete documentation-based assessment, with sources and dates against every claim.
- Alternatives. OpenAI alternatives — organized by the reason you would leave, not as a ranked list.
- Head to head. Cartesia vs OpenAI, OpenAI vs Speechify.
- In our buying guides. Best AI Voice Generators in 2026: 10 Providers Compared, Best AI Dubbing Software 2026: Only Four Vendors Actually Sell It, Best Free AI Voice Generators 2026: What “Free” Actually Buys You, Most Realistic AI Voices: What Every Vendor Charges for the Claim, Best Multilingual AI Voice Generators: Every Count, Recounted, Real-Time AI Voice: Why the Latency Numbers Cannot Be Compared, Best AI Voice Cloning Software 2026: Price, Access and Consent Compared.
- In our comparisons of published data. AI Voice Generator Pricing Comparison: Every Rate Normalized, Best Text-to-Speech APIs: Limits, Transports and Auth Compared.
- In our explainers. Can You Use AI Voices Commercially? What Ten Vendors Actually Say, How AI Voice Cloning Works: What Ten Vendors Require of You, How AI Voice Generators Work, According to the Vendors Themselves.
- By what you are making. AI Voice for Audiobooks: Will Chapter 20 Match Chapter 1?, AI Voice for Business: What Survives a Procurement Review, AI Voice for eLearning: Courses Change, and the Audio Must Follow, AI Voice for Podcasts: Only Three Providers Do Two Speakers, AI Voice for YouTube: What You Must Disclose Before You Publish.
What we could not verify
Honesty about gaps beats a page that looks complete. Standing limits on everything above, as of September 3, 2026:
- How OpenAI sounds. Our listening benchmark is designed but not running. Nothing on this page or on any page it links to ranks OpenAI on audio quality, naturalness or speed.
- Anything OpenAI does not publish. Every page indexed here is built from the vendor’s own documentation. Where that documentation is silent, we record the silence rather than estimating around it.
- Whether a documented feature works as documented. We read documentation; we did not create an account or make an API call.
- Currency. Each linked page carries its own check date. This index inherits them and does not re-verify on load.
Change log
September 3, 2026 — First publication. Assembled from pages already published on this site; no new vendor claim was introduced by this page.
Published September 3, 2026 · This page shows no audio and reports no listening results, because our benchmark has not run. Corrections are recorded with their date on the corrections page. Independence note: this page contains no affiliate links, and no vendor paid for placement or influenced the order — see how we make money and our editorial policy.