mp3 or wav for an AI voiceover? What each allows on Sume TTS
Sume TTS defaults to mp3 at 44.1 kHz. Choose wav for slicing and exact joins, mu-law at 8 kHz for phones, and mp3 for delivery: what each container permits.

Choose wav for anything you will slice, join or caption, and mp3 for the file you finally deliver. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps, and only wav and raw containers can emit sentence audio slices or join without extra padding.
What the request allows
output_format takes a container of mp3, wav or raw, a sample_rate from 8,000 to 48,000 Hz and, for wav and raw, an encoding of pcm_s16le, pcm_f32le, pcm_mulaw or pcm_alaw. Price does not depend on any of it: the bill is per character.
| Choice | Allows | Does not allow | Use for |
|---|---|---|---|
| mp3, 44,100 Hz, 128 kbps (default) | Small files, easy playback | Sentence audio slices; sample-exact joins | Final delivery |
wav, pcm_s16le | Slices with emit_audio; sample-exact concat and split | Small files | Editing, captions, timelines |
| raw | Slices with emit_audio | Direct playback in most players | Pipelines that handle headerless PCM |
wav or raw, pcm_mulaw, 8,000 Hz | Telephony-style audio | Hi-fi playback | Phone lines |
Why mp3 hurts in an editing chain
Sume's timeline audio page says wav is the sample-exact format, and that mp3 adds encoder priming padding at the start and end of each file. Join three mp3 takes and you add padding at every seam, which is audible as a small gap or click. Join the same three takes as wav and the join is in the sample domain, with no re-synthesis and no silence added.
The practical rule is to produce wav at every step and encode to mp3 once, at the end, outside the chain.
A request that supports slicing
Sentence segmentation needs word timestamps, and emit_audio needs wav or raw. A request that asks for audio slices from an mp3 job is rejected, so put the container in the first request, not in a retake.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wav-slices-001" \
-d '{
"transcript": "First line. Second line.",
"voice": { "id": "VOICE_ID" },
"output_format": { "container": "wav", "sample_rate": 44100 },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence", "emit_audio": true }
}'Sample rate and the phone case
For a call-centre style voice, Microsoft's MAI-Voice-2.1 page recommends its Flash model for call centres and voice assistants (read 2026-10-05). If a similar job on Sume feeds a phone line, ask for pcm_mulaw or pcm_alaw at 8,000 Hz so the audio already matches the line, and listen to it on a handset, because a clean 44.1 kHz voice and an 8 kHz voice sound very different.
Remember that Sume's TTS endpoints are asynchronous jobs, with a sync wait of at most 30 seconds. They produce a finished file, not a live stream, so a phone integration still needs your own playback layer.
A short rule set
- Work in wav at 44,100 Hz unless a channel needs less.
- Encode to mp3 once, at delivery.
- Keep one container per job chain; mixed layouts fail concat with
audio_parts_channel_mismatch.
Delivery targets
Ad platforms and social uploaders usually take mp3 or the audio of an mp4, so the final encode is rarely a problem. The risk is the opposite: shipping an mp3 into the editing chain to save disk space, then discovering at the join that every seam has a click. Disk is cheap and a re-render is not, since a Timeline render alone is $0.10 a minute.
If a file must be small in the middle of a chain, such as a long podcast take, keep it as wav in storage and let the pipeline decide when to compress. Name files with their container so nobody has to open one to find out.
What this post cannot tell you
The docs describe the containers and encodings that the request accepts. They do not rank them by sound quality on a given player, so listen to your own output on the devices your audience uses before you set a house default.
Sources
Related posts
More in Developers
- Multi-reference image prompts when the API has no role field
Ideogram's app lets you tag references with @. Sume's input_references holds only a URL. How to say which image is the subject, style or layout in the prompt.
- Multi-turn image edits over Agent Completions: pass the last output
Agent Completions keep no conversation, so each edit turn is a new run. Attach the previous output.images URL as the next input and keep a cap per turn.
- Music API image_url null: clear a reused request body, with Python
Send image_url as null on Sume music requests only to clear an image from a reused request object. A short urllib script posts to the Music Router.
- Nano Banana 2 Lite's 10 aspect ratios: what to send for 4:5
Lite accepts 10 ratios including 4:5 and 21:9. On Sume, Instagram 1080x1350 is aspect_ratio 4:5, and Nano Banana Pro or 2 list it; Imagen and Grok do not.
Written by Sume