WAV or MP3 for AI narration you will trim or join: Sume defaults
Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps, but Timeline audio says mp3 re-adds priming padding at every edge. Use wav for narration you cut or join.

Request wav, not mp3, for any Sume narration you will trim, split or join. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps, which is fine for a file you only play. But the Timeline audio docs say mp3 output re-adds priming padding at every edge and recommend wav, which is pcm_s16le and sample-exact, whenever a file will be joined again or drives lip-sync. Per-sentence slices also need wav or raw.
The defaults and the slice rule come from the Sume API reference, and the padding note and output limits from Timeline audio and Timeline 1.0, read on 2026-10-03. I did not measure the padding in milliseconds, so I do not give one.
What changes between the two?
Mostly what you can do next. File size is secondary: it is simple arithmetic from the format.
| Question | wav | mp3 |
|---|---|---|
| TTS default | Pass it explicitly | Yes, 44,100 Hz, 128 kbps |
Per-sentence slices (emit_audio) | Yes, wav or raw | Timings only, no slices |
| Join or split edges | Sample-exact | Priming padding added at every edge |
| Size of 60 seconds, one channel | About 5,292 KB at 44.1 kHz 16-bit | About 960 KB at 128 kbps |
How do I ask for wav?
Set output_format to a wav container with pcm_s16le and 44,100 Hz, the same recipe Sume's own tool descriptions give for a voice spine. This script prints the body and the size of a take at different lengths.
RATE = 0.0475 / 1000
def wav_bytes(seconds, rate=44100, channels=1):
return seconds * rate * 2 * channels # pcm_s16le is 2 bytes per sample
def mp3_bytes(seconds, bit_rate=128000):
return seconds * bit_rate // 8
for s in (30, 60, 180):
print(s, "s: wav", wav_bytes(s) // 1000, "KB per channel; mp3", mp3_bytes(s) // 1000, "KB")What about lip-sync and avatars?
The same rule holds there. Timeline audio's docs say to keep wav when a file drives lip-sync, and the TTS reference says to pass wav, pcm_s16le and 44,100 Hz explicitly when the audio feeds an avatar mux. So one wav take serves the join, the split and the avatar clip, and you convert once at the end if you need a smaller file.
When is mp3 still the right choice?
When the file is final and goes straight to a listener, mp3 is five to six times smaller. If you must deliver mp3 after editing, cut and join in wav, then request mp3 as the last step with Timeline audio's output.format, so the padding is added once. A join or split is one $0.01 job whatever the format, and the produced audio is capped at 1,800 seconds.
Sources
Related posts
More in Developers
- Sume webhook handler in Node: store first, answer 2xx, process later
Sume gives a webhook endpoint 10 seconds per attempt. Verify the raw body, persist the event, answer 2xx, then do slow work off the request path. Node sample.
- Empty webhook secret in staging: fail closed at boot, not per request
An unset SUME_COM_WEBHOOK_SIGNING_SECRET makes an HMAC over an empty key that anyone can forge. Refuse to start, then accept either entry during a rotation.
- Test a webhook endpoint before go-live: a Sume CI gate (Python)
Use POST /v1/webhooks/test-deliveries to fire a signed webhook.test at your deployed URL and fail the deploy unless it answers 2xx. Python script included.
- Webhook timestamp tolerance: Sume 300 s versus Stripe 5 min
Sume's verifyWebhook rejects deliveries older than 300 seconds by default. Why clock skew matters, why toleranceSeconds 0 is a trap, and how to fix drift.
Written by Sume