mp3 or wav for an AI voiceover? What each allows on Sume TTS

Sume TTS defaults to mp3 at 44.1 kHz. Choose wav for slicing and exact joins, mu-law at 8 kHz for phones, and mp3 for delivery: what each container permits.

5 min readSume
All posts

Choose wav for anything you will slice, join or caption, and mp3 for the file you finally deliver. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps, and only wav and raw containers can emit sentence audio slices or join without extra padding.

What the request allows

output_format takes a container of mp3, wav or raw, a sample_rate from 8,000 to 48,000 Hz and, for wav and raw, an encoding of pcm_s16le, pcm_f32le, pcm_mulaw or pcm_alaw. Price does not depend on any of it: the bill is per character.

Container choices on Sume TTS and what they allow (Sume API reference and Timeline audio page, read 2026-10-05)
ChoiceAllowsDoes not allowUse for
mp3, 44,100 Hz, 128 kbps (default)Small files, easy playbackSentence audio slices; sample-exact joinsFinal delivery
wav, pcm_s16leSlices with emit_audio; sample-exact concat and splitSmall filesEditing, captions, timelines
rawSlices with emit_audioDirect playback in most playersPipelines that handle headerless PCM
wav or raw, pcm_mulaw, 8,000 HzTelephony-style audioHi-fi playbackPhone lines

Why mp3 hurts in an editing chain

Sume's timeline audio page says wav is the sample-exact format, and that mp3 adds encoder priming padding at the start and end of each file. Join three mp3 takes and you add padding at every seam, which is audible as a small gap or click. Join the same three takes as wav and the join is in the sample domain, with no re-synthesis and no silence added.

The practical rule is to produce wav at every step and encode to mp3 once, at the end, outside the chain.

A request that supports slicing

Sentence segmentation needs word timestamps, and emit_audio needs wav or raw. A request that asks for audio slices from an mp3 job is rejected, so put the container in the first request, not in a retake.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wav-slices-001" \
  -d '{
    "transcript": "First line. Second line.",
    "voice": { "id": "VOICE_ID" },
    "output_format": { "container": "wav", "sample_rate": 44100 },
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence", "emit_audio": true }
  }'

Sample rate and the phone case

For a call-centre style voice, Microsoft's MAI-Voice-2.1 page recommends its Flash model for call centres and voice assistants (read 2026-10-05). If a similar job on Sume feeds a phone line, ask for pcm_mulaw or pcm_alaw at 8,000 Hz so the audio already matches the line, and listen to it on a handset, because a clean 44.1 kHz voice and an 8 kHz voice sound very different.

Remember that Sume's TTS endpoints are asynchronous jobs, with a sync wait of at most 30 seconds. They produce a finished file, not a live stream, so a phone integration still needs your own playback layer.

A short rule set

  • Work in wav at 44,100 Hz unless a channel needs less.
  • Encode to mp3 once, at delivery.
  • Keep one container per job chain; mixed layouts fail concat with audio_parts_channel_mismatch.

Delivery targets

Ad platforms and social uploaders usually take mp3 or the audio of an mp4, so the final encode is rarely a problem. The risk is the opposite: shipping an mp3 into the editing chain to save disk space, then discovering at the join that every seam has a click. Disk is cheap and a re-render is not, since a Timeline render alone is $0.10 a minute.

If a file must be small in the middle of a chain, such as a long podcast take, keep it as wav in storage and let the pipeline decide when to compress. Name files with their container so nobody has to open one to find out.

What this post cannot tell you

The docs describe the containers and encodings that the request accepts. They do not rank them by sound quality on a given player, so listen to your own output on the devices your audience uses before you set a house default.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume