TTS WAV or MP3 for video: keep WAV until the last mux
Sume TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps. If the audio will be joined, sliced or muxed into video, ask for wav; mp3 adds padding at every edge.
For TTS that ends up in a video, ask for wav and keep it until the final step. Sume TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps, and the API notes say to pass wav with pcm_s16le at 44100 explicitly when the audio feeds the avatar mux. The reason is edges: an mp3 re-introduces encoder padding every time it is cut or joined, while wav stays sample-exact.
The facts below come from the TTS request schema in the OpenAPI document behind the API reference and the Timeline audio page, read 2026-09-29.
What is the default if I send nothing?
output_format defaults to an mp3 container with sample_rate 44100 and bit_rate 128000. That is fine for a file you play once and never touch again. It is the wrong default for a file that goes on a timeline, gets sliced by sentence, or is muxed under video.
Why does mp3 hurt when I join or cut?
The Timeline audio schema describes wav (pcm_s16le, default) as staying sample-exact, so it can be joined again or used to drive Avatar 1.0 image-to-video without picking up encoder delay. It describes mp3 as smaller but re-introducing priming padding at every edge. The Timeline audio page adds a rule of thumb: keep wav when the file will be joined again or drives lip-sync.
| Next step | Container | Why |
|---|---|---|
| Join with Timeline audio concat | wav | Sample-exact, no padding at the seams |
Sentence slices with audio_url | wav or raw | Segment audio is only sliced for wav and raw |
| Drive a lip-sync or avatar clip | wav | No encoder delay at the start |
| One-off listen or download | mp3 | Smaller file, nothing to cut later |
What exactly should I send?
Set the container, the sample rate, and the PCM encoding together. The schema accepts wav or raw containers with PCM encodings such as pcm_s16le, and sample rates from 8000 to 48000 Hz.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Meet the new travel mug.",
"avatar_handle": "@your-avatar",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" }
}'Does the same rule apply to audio I pull out of a video?
Yes. The Audio detach API extracts a video's audio track and defaults to wav, described as what timeline_create audio.url, timeline_audio and the avatar mux want, with mp3 as the option. So the pipeline stays wav from source to the last join, and mp3 belongs only at the very end if file size matters. In practice that means one rule for the whole chain: request wav from TTS, keep wav through any slicing, joining or detaching, and choose a container for the deliverable only when nothing downstream will touch the audio again. If you are unsure whether a file will be reused, keep the wav; converting down later is a choice you can still make, while the padding from an early mp3 cannot be undone.
Sources
Related posts
More in Developers
- Let the API pick the video model: sume/auto for vertical UGC clips
Send model sume/auto to POST /v1/videos and Sume picks the family. The response echoes sume/auto and never names the model. When to pin a model instead.
- 400 unknown_parameter on a Format run: read the suggestion
Sume rejects a top-level field it does not know with 400 unknown_parameter and names the likely one. Why webook_url and thread_id fail, and the fix.
- Veo 3.1 personGeneration: allow_adult vs allow_all by mode
Veo 3.1 personGeneration is allow_all for text-to-video, allow_adult for image modes, and allow_adult only in some regions. Sume has no such field.
- verifyWebhook async: forget await and you get a promise
verifyWebhook in @sume-com/sdk is async. Without await, a Promise is truthy, so your signature check never rejects. The await, raw body and replay rules.
Written by Sume