ElevenLabs Opus 48 kHz output vs Sume TTS mp3, wav and raw formats
ElevenLabs lists Opus, MP3, PCM and mu-law outputs. Sume TTS 1.0 offers mp3, wav and raw with six sample rates. See which fits your pipeline.

What ElevenLabs lists
The ElevenLabs text to speech page lists these output families: MP3 from 22.05 to 44.1 kHz at 32 to 192 kbps, PCM S16LE from 16 to 44.1 kHz, mu-law and A-law at 8 kHz, and Opus at 48 kHz with 32 to 192 kbps.
If your product plays audio through an Opus-native stack, getting Opus directly saves a conversion step.
What Sume TTS 1.0 returns
Sume TTS 1.0 takes an output_format with a container of mp3, wav or raw. The default is mp3 at 44.1 kHz and 128 kbps. Sample rate can be 8000, 16000, 22050, 24000, 44100 or 48000. MP3 bit rate runs from 32k to 192k. For wav and raw you choose a PCM encoding: pcm_f32le, pcm_s16le, pcm_mulaw or pcm_alaw.
| Need | ElevenLabs | Sume TTS 1.0 |
|---|---|---|
| Telephony 8 kHz mu-law | Yes, mu-law and A-law at 8 kHz | Yes, pcm_mulaw and pcm_alaw at 8000 Hz |
| MP3 | 22.05-44.1 kHz, 32-192 kbps | 8000-48000 Hz, 32k-192k |
| Lossless PCM 16-bit | PCM S16LE 16-44.1 kHz | pcm_s16le in wav or raw |
| Opus 48 kHz | Yes, 32-192 kbps | Not listed; convert from wav |
Getting to Opus from Sume
If you need Opus, ask Sume for 48000 Hz wav and encode with your own tool. Starting from lossless means one lossy step, not two. Do not render to mp3 first and then re-encode, because each lossy generation adds artifacts.
- Set
output_formatto containerwav,sample_rate48000 and encodingpcm_s16le. - Encode to Opus in your build step.
- Keep the wav if you still need to join clips, then encode last.
Choosing between them
Pick on the pipeline, not the codec name. A phone system wants 8 kHz mu-law, which both support. A video pipeline wants wav so Timeline audio can join cleanly and a lip-sync model gets lossless input. A web page that just plays a clip is happy with the default mp3.
Which sample rate should you pick
Match the destination. Telephony wants 8000 Hz; speech models for transcription often use 16000 Hz; video work wants 48000 Hz, since that is the rate video containers usually carry. Rendering at the destination rate avoids a resample step.
Raw output is headerless PCM, which suits a pipeline that already knows the rate and encoding. A wav carries a header, which makes it self-describing and safer to hand to other tools.
Takeaway
Sume covers mp3, wav and raw PCM including 8 kHz telephony encodings. Opus is the one ElevenLabs output Sume does not list, and wav at 48 kHz is the clean source for it.
Sources
Related posts
More in Comparisons
- ElevenLabs previous_text and next_text vs Sume TTS: split scripts
ElevenLabs lets you pass previous_text and next_text for prosody continuity. Sume TTS has no such fields, so here is how to keep long scripts smooth.
- ElevenLabs TTS seed and free regenerations vs Sume TTS retry cost
ElevenLabs offers a seed and up to two free regenerations of identical requests. On Sume a retry with the same Idempotency-Key returns the same job. Compare.
- ElevenLabs v4 Turbo vs Flash vs v3: price per 1,000 characters
ElevenLabs API list per 1K characters: v4 Turbo $0.011 until Oct 12 (then $0.04), Flash/Turbo $0.04, v3 Multilingual $0.08, v4 $0.022 (then $0.08).
- Face and body swap video AI: Recast vs Sume's avatar face swap
Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.
Written by Sume