ElevenLabs TTS output_format: 192 kbps needs Creator, PCM needs Pro
The ElevenLabs text to speech reference ties 192 kbps MP3 to Creator and PCM or WAV to Pro. A format table, a fallback chooser in Python, and the seed range.

On the ElevenLabs text to speech endpoint, 192 kbps MP3 needs the Creator tier or higher, and PCM and WAV output need Pro or higher.
If your code asks for a format the account cannot use, it fails at request time, so pick the format from the plan or fall back deliberately.
What the reference lists
From the API reference page read 2026-10-03.
| Item | Value |
|---|---|
| Formats | MP3, Opus, PCM, WAV, mu-law |
| 192 kbps MP3 | Creator tier or above |
| PCM and WAV | Pro tier or above |
seed range | 0 to 4,294,967,295 |
| Context request ids | Maximum 3 request_ids |
A format chooser with a fallback
The chooser takes the plan name and what you want and returns an output_format string. The exact string values (codec, rate and bitrate combinations) come from the reference's enum, so the script uses placeholders you must replace with the values listed there rather than guessing them.
ORDER = ["free", "starter", "creator", "pro"]
# Fill these from the output_format enum on the reference page.
FORMATS = {
"mp3_192": {"min_plan": "creator", "value": "<mp3 192 kbps value>"},
"pcm": {"min_plan": "pro", "value": "<pcm value>"},
"wav": {"min_plan": "pro", "value": "<wav value>"},
"mp3_default": {"min_plan": "free", "value": "<default mp3 value>"},
}
def choose(plan, wanted):
plan = plan.lower()
if plan not in ORDER:
raise ValueError(f"unknown plan: {plan}")
for key in wanted + ["mp3_default"]:
need = FORMATS[key]["min_plan"]
if ORDER.index(plan) >= ORDER.index(need):
return key, FORMATS[key]["value"]
raise RuntimeError("no usable format")
print(choose("creator", ["wav", "mp3_192"]))Practical notes
Choices that follow from the table.
- If a later step needs lossless audio, such as joining or splitting, ask for PCM or WAV only if the plan allows it. Otherwise decode the MP3 once and keep it as the working copy.
- Store the
seedwith each take. The reference gives the valid range, so a stored value lets you ask for the same take again. request_idsfor continuity is capped at 3. Keep a rolling window of the last three rather than every previous id.- Check the plan name against what the account actually has, not what the pricing page says it should have. The chooser above only knows the names you give it.
Joining takes afterward
If you chain several voice takes into one soundtrack on Sume, the timeline audio docs list a $0.01 concat or split job with a WAV default, and note that MP3 output carries encoder priming padding. Source format therefore decides whether a join is sample exact, which is one more reason to take the lossless format when the plan permits.
Sources
Related posts
More in Developers
- Gemini CLI v0.63 plan execution in CI: gate paid Sume calls first
Gemini CLI preview v0.63.0 adds autonomous plan execution in non-interactive mode. Before unattended runs, gate Sume paid tools with dry_run and max_spend_usd.
- Google's June 15 deprecation notice gave 15 and 63 days: run a drill
Google announced Veo and Imagen 4 deprecations on Jun 15, 2026 with shutdowns Jun 30 and Aug 17. Here is a five-step drill that fits inside the shorter window.
- Grok Imagine video API: 15 s, 5 references, request-ID polling
xAI's video guide for grok-imagine-video-1.5 lists up to 15 seconds, up to 5 reference images and async polling by request ID. The same loop on Sume jobs.
- How long AI video vendors keep your file: Veo 2 days, Higgsfield 7+
Veo keeps videos two days, Higgsfield files at least seven, Sora Batch outputs were kept 24 hours. Retention facts and a download-on-complete script.
Written by Sume