OpenAI TTS: 13 built-in voices vs 9 legacy ones, pick the model first
OpenAI's guide lists gpt-4o-mini-tts with 13 built-in voices and legacy tts-1 and tts-1-hd with 9. Formats, custom voice consent and disclosure, as a checklist.

In OpenAI's text to speech guide, gpt-4o-mini-tts has 13 built-in voices, while the legacy tts-1 and tts-1-hd have 9. Choose the model first, because the voice list depends on it.
The guide also requires disclosing that a voice is AI-generated, and approved custom voices need a consent recording.
Facts from the guide
As read on 2026-10-03.
| Item | Value |
|---|---|
| Current model | gpt-4o-mini-tts, 13 built-in voices |
| Legacy models | tts-1 and tts-1-hd, 9 voices |
| Output formats | mp3, opus, aac, flac, wav, pcm |
| Custom voices | Approved voices need a consent recording |
| Disclosure | AI-generated voice must be disclosed to listeners |
Guard the model and voice pair
A request that pairs a model with a voice it does not offer fails at the API, which is a poor way to find out during a batch. This small check keeps the pairing in your own config. The voice names are yours to fill from the guide's lists; the script does not guess them.
VOICES = {
"gpt-4o-mini-tts": set(), # fill with the 13 names from the guide
"tts-1": set(), # fill with the 9 legacy names
"tts-1-hd": set(), # same 9
}
FORMATS = {"mp3", "opus", "aac", "flac", "wav", "pcm"}
def check(model, voice, fmt):
if model not in VOICES:
raise ValueError(f"unknown model {model}")
if not VOICES[model]:
raise RuntimeError(f"fill VOICES[{model!r}] from the guide first")
if voice not in VOICES[model]:
raise ValueError(f"{voice} is not a {model} voice")
if fmt not in FORMATS:
raise ValueError(f"unsupported format {fmt}")
# check("gpt-4o-mini-tts", "<voice>", "wav")Rules around voices
Policy items that belong in a delivery checklist, not in a prompt.
- Disclose AI-generated voice where the guide requires it. Put the disclosure in the deliverable's metadata or credits, not only in a ticket.
- For custom voices, store the consent recording with the voice id. The guide ties approval to that recording.
- Pick the output format for the next step. WAV or PCM suits editing; MP3 or Opus suits delivery.
- Pin the model id in config. A legacy model id and a current one can sound different with the same voice name.
Assembling the finished video
A narration file is one input to a finished clip. If you then need to place it against music or cut it into ranges on Sume, the timeline audio docs describe concat and split jobs at $0.01 each with a WAV default, so a WAV narration take keeps those edits clean.
Sources
Related posts
More in Developers
- Pause between narration lines: TTS pause markers or audio concat?
Deepgram Flux TTS allows pause markers of 500 to 3000 ms, max 8 per request. Sume's timeline audio concat joins lines with no gap. Where to put the pause.
- Performance Max text limits: 30, 90, 90 and 25 characters, counted
Performance Max headlines run 30 characters, long headlines 90, descriptions 90, business name 25. A short script counts generated copy before upload.
- Push or poll for a finished render: listen, webhook or jobs_wait
MCP 2026-07-28 adds subscriptions/listen. For a render that takes minutes, compare a listen stream, a signed webhook and jobs_wait, with a Python verifier.
- Python 3.10 is end of life: a stdlib Sume webhook verifier
Python 3.10 has reached end of life. A standard-library verifier for Sume's signed webhooks that refuses an empty secret and accepts rotated signatures.
Written by Sume