TTS WAV or MP3 for video: keep WAV until the last mux

Sume TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps. If the audio will be joined, sliced or muxed into video, ask for wav; mp3 adds padding at every edge.

4 min readSume
All posts

For TTS that ends up in a video, ask for wav and keep it until the final step. Sume TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps, and the API notes say to pass wav with pcm_s16le at 44100 explicitly when the audio feeds the avatar mux. The reason is edges: an mp3 re-introduces encoder padding every time it is cut or joined, while wav stays sample-exact.

The facts below come from the TTS request schema in the OpenAPI document behind the API reference and the Timeline audio page, read 2026-09-29.

What is the default if I send nothing?

output_format defaults to an mp3 container with sample_rate 44100 and bit_rate 128000. That is fine for a file you play once and never touch again. It is the wrong default for a file that goes on a timeline, gets sliced by sentence, or is muxed under video.

Why does mp3 hurt when I join or cut?

The Timeline audio schema describes wav (pcm_s16le, default) as staying sample-exact, so it can be joined again or used to drive Avatar 1.0 image-to-video without picking up encoder delay. It describes mp3 as smaller but re-introducing priming padding at every edge. The Timeline audio page adds a rule of thumb: keep wav when the file will be joined again or drives lip-sync.

Which container to request from TTS 1.0 by next step, read 2026-09-29.
Next stepContainerWhy
Join with Timeline audio concatwavSample-exact, no padding at the seams
Sentence slices with audio_urlwav or rawSegment audio is only sliced for wav and raw
Drive a lip-sync or avatar clipwavNo encoder delay at the start
One-off listen or downloadmp3Smaller file, nothing to cut later

What exactly should I send?

Set the container, the sample rate, and the PCM encoding together. The schema accepts wav or raw containers with PCM encodings such as pcm_s16le, and sample rates from 8000 to 48000 Hz.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "Meet the new travel mug.",
    "avatar_handle": "@your-avatar",
    "output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" }
  }'

Does the same rule apply to audio I pull out of a video?

Yes. The Audio detach API extracts a video's audio track and defaults to wav, described as what timeline_create audio.url, timeline_audio and the avatar mux want, with mp3 as the option. So the pipeline stays wav from source to the last join, and mp3 belongs only at the very end if file size matters. In practice that means one rule for the whole chain: request wav from TTS, keep wav through any slicing, joining or detaching, and choose a container for the deliverable only when nothing downstream will touch the audio again. If you are unsure whether a file will be reused, keep the wav; converting down later is a choice you can still make, while the padding from an early mp3 cannot be undone.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume