ElevenLabs TTS output_format: 192 kbps needs Creator, PCM needs Pro

The ElevenLabs text to speech reference ties 192 kbps MP3 to Creator and PCM or WAV to Pro. A format table, a fallback chooser in Python, and the seed range.

4 min readSume
All posts

On the ElevenLabs text to speech endpoint, 192 kbps MP3 needs the Creator tier or higher, and PCM and WAV output need Pro or higher.

If your code asks for a format the account cannot use, it fails at request time, so pick the format from the plan or fall back deliberately.

What the reference lists

From the API reference page read 2026-10-03.

ElevenLabs TTS output facts (read 2026-10-03)
ItemValue
FormatsMP3, Opus, PCM, WAV, mu-law
192 kbps MP3Creator tier or above
PCM and WAVPro tier or above
seed range0 to 4,294,967,295
Context request idsMaximum 3 request_ids

A format chooser with a fallback

The chooser takes the plan name and what you want and returns an output_format string. The exact string values (codec, rate and bitrate combinations) come from the reference's enum, so the script uses placeholders you must replace with the values listed there rather than guessing them.

ORDER = ["free", "starter", "creator", "pro"]

# Fill these from the output_format enum on the reference page.
FORMATS = {
    "mp3_192": {"min_plan": "creator", "value": "<mp3 192 kbps value>"},
    "pcm": {"min_plan": "pro", "value": "<pcm value>"},
    "wav": {"min_plan": "pro", "value": "<wav value>"},
    "mp3_default": {"min_plan": "free", "value": "<default mp3 value>"},
}

def choose(plan, wanted):
    plan = plan.lower()
    if plan not in ORDER:
        raise ValueError(f"unknown plan: {plan}")
    for key in wanted + ["mp3_default"]:
        need = FORMATS[key]["min_plan"]
        if ORDER.index(plan) >= ORDER.index(need):
            return key, FORMATS[key]["value"]
    raise RuntimeError("no usable format")

print(choose("creator", ["wav", "mp3_192"]))

Practical notes

Choices that follow from the table.

  • If a later step needs lossless audio, such as joining or splitting, ask for PCM or WAV only if the plan allows it. Otherwise decode the MP3 once and keep it as the working copy.
  • Store the seed with each take. The reference gives the valid range, so a stored value lets you ask for the same take again.
  • request_ids for continuity is capped at 3. Keep a rolling window of the last three rather than every previous id.
  • Check the plan name against what the account actually has, not what the pricing page says it should have. The chooser above only knows the names you give it.

Joining takes afterward

If you chain several voice takes into one soundtrack on Sume, the timeline audio docs list a $0.01 concat or split job with a WAV default, and note that MP3 output carries encoder priming padding. Source format therefore decides whether a join is sample exact, which is one more reason to take the lossless format when the plan permits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume