OpenAI TTS: 13 built-in voices vs 9 legacy ones, pick the model first

OpenAI's guide lists gpt-4o-mini-tts with 13 built-in voices and legacy tts-1 and tts-1-hd with 9. Formats, custom voice consent and disclosure, as a checklist.

4 min readSume
All posts

In OpenAI's text to speech guide, gpt-4o-mini-tts has 13 built-in voices, while the legacy tts-1 and tts-1-hd have 9. Choose the model first, because the voice list depends on it.

The guide also requires disclosing that a voice is AI-generated, and approved custom voices need a consent recording.

Facts from the guide

As read on 2026-10-03.

OpenAI TTS facts (read 2026-10-03)
ItemValue
Current modelgpt-4o-mini-tts, 13 built-in voices
Legacy modelstts-1 and tts-1-hd, 9 voices
Output formatsmp3, opus, aac, flac, wav, pcm
Custom voicesApproved voices need a consent recording
DisclosureAI-generated voice must be disclosed to listeners

Guard the model and voice pair

A request that pairs a model with a voice it does not offer fails at the API, which is a poor way to find out during a batch. This small check keeps the pairing in your own config. The voice names are yours to fill from the guide's lists; the script does not guess them.

VOICES = {
    "gpt-4o-mini-tts": set(),   # fill with the 13 names from the guide
    "tts-1": set(),             # fill with the 9 legacy names
    "tts-1-hd": set(),          # same 9
}
FORMATS = {"mp3", "opus", "aac", "flac", "wav", "pcm"}

def check(model, voice, fmt):
    if model not in VOICES:
        raise ValueError(f"unknown model {model}")
    if not VOICES[model]:
        raise RuntimeError(f"fill VOICES[{model!r}] from the guide first")
    if voice not in VOICES[model]:
        raise ValueError(f"{voice} is not a {model} voice")
    if fmt not in FORMATS:
        raise ValueError(f"unsupported format {fmt}")

# check("gpt-4o-mini-tts", "<voice>", "wav")

Rules around voices

Policy items that belong in a delivery checklist, not in a prompt.

  • Disclose AI-generated voice where the guide requires it. Put the disclosure in the deliverable's metadata or credits, not only in a ticket.
  • For custom voices, store the consent recording with the voice id. The guide ties approval to that recording.
  • Pick the output format for the next step. WAV or PCM suits editing; MP3 or Opus suits delivery.
  • Pin the model id in config. A legacy model id and a current one can sound different with the same voice name.

Assembling the finished video

A narration file is one input to a finished clip. If you then need to place it against music or cut it into ranges on Sume, the timeline audio docs describe concat and split jobs at $0.01 each with a WAV default, so a WAV narration take keeps those edits clean.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume