ElevenLabs TTS output_format list vs Sume container and encoding

ElevenLabs names output_format strings like opus_48000_64 and ulaw_8000. Sume splits it into container, sample rate, bit rate and encoding, with no Opus.

5 min readSume
All posts

ElevenLabs picks an audio format with one string, output_format, such as mp3_44100_128, opus_48000_64 or ulaw_8000. Sume's TTS takes an object instead, output_format with a container (mp3, wav or raw), a sample_rate, an MP3 bit_rate and a PCM encoding. The practical gap: Sume does not produce Opus or A-law/mu-law as a named container string, and it does produce mu-law and A-law as PCM encodings inside wav or raw.

The ElevenLabs side is from its Text to Speech convert reference, read on 2026-10-02. The Sume side is the TTS request schema in the live OpenAPI linked from the API reference.

Which formats does the ElevenLabs reference list?

The convert page lists MP3 formats mp3_22050_32, mp3_24000_48 and mp3_44100_32, _64, _96, _128 and _192; PCM formats from pcm_8000 up to pcm_48000; WAV formats from wav_8000 up to wav_48000; Opus formats opus_48000_32 through opus_48000_192; and the legacy alaw_8000 and ulaw_8000.

Each string fixes the codec, the sample rate and, for the lossy ones, the bit rate in one token. That is compact, and it makes the list long.

How does Sume express the same choice?

Sume's default is MP3 at 44,100 Hz and 128 kbps. Override it with an object. container is mp3, wav or raw. sample_rate is one of 8000, 16000, 22050, 24000, 44100 or 48000. bit_rate is 32000, 64000, 96000, 128000 or 192000 and is required for MP3 when you override the defaults. encoding is pcm_f32le, pcm_s16le, pcm_mulaw or pcm_alaw and applies to wav and raw.

The request that gets a phone-call test file looks like this:

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-8khz-mulaw-001" \
  -d '{
    "transcript": "Thanks for calling. Press one for sales.",
    "avatar_handle": "speaker",
    "output_format": {
      "container": "wav",
      "sample_rate": 8000,
      "encoding": "pcm_mulaw"
    }
  }'

Where do the two vocabularies line up?

ElevenLabs strings from the convert reference; Sume fields from the TTS request schema, read 2026-10-02.
You wantElevenLabs stringSume object
MP3, 44.1 kHz, 128 kbpsmp3_44100_128mp3, 44100, 128000 (the default)
Uncompressed 48 kHzwav_48000 or pcm_48000wav or raw, 48000, pcm_s16le
Telephony mu-law, 8 kHzulaw_8000wav, 8000, pcm_mulaw
Telephony A-law, 8 kHzalaw_8000wav, 8000, pcm_alaw
Opus, 48 kHzopus_48000_64Not offered
MP3 at 22.05 kHz, 32 kbpsmp3_22050_32mp3, 22050, 32000

When does the choice change the rest of the pipeline?

Two Sume rules matter more than the codec. First, sentence segmentation needs timestamps.words: true, and it only returns per-sentence audio slices (audio_url) when the container is wav or raw; with MP3 you get timings but no slice files. Second, the Sume docs recommend an explicit WAV with pcm_s16le at 44,100 Hz when the audio will be muxed into an avatar video, rather than the MP3 default.

If you need Opus, for a WebRTC leg or a chat app, produce WAV on Sume and encode Opus in your own step. Sume's media tools do not list Opus as an output.

How do you pick a format for a real job?

Start from where the file goes next, not from the codec. For a voiceover that a human will review and then drop into an edit, keep the MP3 default or choose a 192 kbps MP3; it is small and plays everywhere. For anything that will be cut, joined or lip-synced, ask for WAV so no encoder padding is added at the seams. Sume's own timeline audio docs make the same point for joins: WAV is sample-exact, while MP3 re-adds priming padding at every edge.

For test calls into a phone system, match the carrier format. Mu-law at 8,000 Hz is the usual North American telephony rate and A-law the usual European one, and both are available as PCM encodings. Generate a short line first and confirm that your IVR plays it before you render a long prompt set. A failed format choice costs a re-run, so keep the first test to one sentence.

What does Sume not do?

Check any combination you depend on against the live OpenAPI document, because Markdown tables can lag it.

  • It does not accept ElevenLabs' string names; send the object form.
  • It does not return Opus, and the docs list no Opus output on the TTS job.
  • It does not stream audio while synthesizing; you read a finished artifact from the job result.
  • It does not publish a latency figure for any container choice.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume