ElevenLabs TTS output_format list vs Sume container and encoding
ElevenLabs names output_format strings like opus_48000_64 and ulaw_8000. Sume splits it into container, sample rate, bit rate and encoding, with no Opus.

ElevenLabs picks an audio format with one string, output_format, such as mp3_44100_128, opus_48000_64 or ulaw_8000. Sume's TTS takes an object instead, output_format with a container (mp3, wav or raw), a sample_rate, an MP3 bit_rate and a PCM encoding. The practical gap: Sume does not produce Opus or A-law/mu-law as a named container string, and it does produce mu-law and A-law as PCM encodings inside wav or raw.
The ElevenLabs side is from its Text to Speech convert reference, read on 2026-10-02. The Sume side is the TTS request schema in the live OpenAPI linked from the API reference.
Which formats does the ElevenLabs reference list?
The convert page lists MP3 formats mp3_22050_32, mp3_24000_48 and mp3_44100_32, _64, _96, _128 and _192; PCM formats from pcm_8000 up to pcm_48000; WAV formats from wav_8000 up to wav_48000; Opus formats opus_48000_32 through opus_48000_192; and the legacy alaw_8000 and ulaw_8000.
Each string fixes the codec, the sample rate and, for the lossy ones, the bit rate in one token. That is compact, and it makes the list long.
How does Sume express the same choice?
Sume's default is MP3 at 44,100 Hz and 128 kbps. Override it with an object. container is mp3, wav or raw. sample_rate is one of 8000, 16000, 22050, 24000, 44100 or 48000. bit_rate is 32000, 64000, 96000, 128000 or 192000 and is required for MP3 when you override the defaults. encoding is pcm_f32le, pcm_s16le, pcm_mulaw or pcm_alaw and applies to wav and raw.
The request that gets a phone-call test file looks like this:
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-8khz-mulaw-001" \
-d '{
"transcript": "Thanks for calling. Press one for sales.",
"avatar_handle": "speaker",
"output_format": {
"container": "wav",
"sample_rate": 8000,
"encoding": "pcm_mulaw"
}
}'Where do the two vocabularies line up?
| You want | ElevenLabs string | Sume object |
|---|---|---|
| MP3, 44.1 kHz, 128 kbps | mp3_44100_128 | mp3, 44100, 128000 (the default) |
| Uncompressed 48 kHz | wav_48000 or pcm_48000 | wav or raw, 48000, pcm_s16le |
| Telephony mu-law, 8 kHz | ulaw_8000 | wav, 8000, pcm_mulaw |
| Telephony A-law, 8 kHz | alaw_8000 | wav, 8000, pcm_alaw |
| Opus, 48 kHz | opus_48000_64 | Not offered |
| MP3 at 22.05 kHz, 32 kbps | mp3_22050_32 | mp3, 22050, 32000 |
When does the choice change the rest of the pipeline?
Two Sume rules matter more than the codec. First, sentence segmentation needs timestamps.words: true, and it only returns per-sentence audio slices (audio_url) when the container is wav or raw; with MP3 you get timings but no slice files. Second, the Sume docs recommend an explicit WAV with pcm_s16le at 44,100 Hz when the audio will be muxed into an avatar video, rather than the MP3 default.
If you need Opus, for a WebRTC leg or a chat app, produce WAV on Sume and encode Opus in your own step. Sume's media tools do not list Opus as an output.
How do you pick a format for a real job?
Start from where the file goes next, not from the codec. For a voiceover that a human will review and then drop into an edit, keep the MP3 default or choose a 192 kbps MP3; it is small and plays everywhere. For anything that will be cut, joined or lip-synced, ask for WAV so no encoder padding is added at the seams. Sume's own timeline audio docs make the same point for joins: WAV is sample-exact, while MP3 re-adds priming padding at every edge.
For test calls into a phone system, match the carrier format. Mu-law at 8,000 Hz is the usual North American telephony rate and A-law the usual European one, and both are available as PCM encodings. Generate a short line first and confirm that your IVR plays it before you render a long prompt set. A failed format choice costs a re-run, so keep the first test to one sentence.
What does Sume not do?
Check any combination you depend on against the live OpenAPI document, because Markdown tables can lag it.
- It does not accept ElevenLabs' string names; send the object form.
- It does not return Opus, and the docs list no Opus output on the TTS job.
- It does not stream audio while synthesizing; you read a finished artifact from the job result.
- It does not publish a latency figure for any container choice.
Sources
Related posts
More in Comparisons
- ElevenLabs Professional Voice Clone: own voice only, 24 h retry
ElevenLabs Professional Voice Cloning only accepts your own voice, needs 30 minutes of audio and a verification step. What Sume's Voices library asks instead.
- ElevenLabs use_pvc_as_ivc: what it does and Sume's voice selector
ElevenLabs' use_pvc_as_ivc flag swaps a professional voice clone for its instant version. Sume's TTS has no such switch: you pick an avatar or a voice id.
- ElevenLabs stability 0.5 and similarity 0.75 vs Sume generation_config
ElevenLabs defaults stability to 0.5 and similarity to 0.75. Sume has no such sliders: generation_config takes only volume, speed and an emotion string.
- fal Agent (early access) vs Sume Agent Completions and hosted MCP
fal's August 2026 agent is early access, with API, CLI and MCP. Sume's agent has Agent Completions over the API and a hosted MCP. What each documents today.
Written by Sume