WAV or MP3 for AI narration you will trim or join: Sume defaults

Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps, but Timeline audio says mp3 re-adds priming padding at every edge. Use wav for narration you cut or join.

4 min readSume
All posts

Request wav, not mp3, for any Sume narration you will trim, split or join. Sume TTS defaults to mp3 at 44,100 Hz and 128 kbps, which is fine for a file you only play. But the Timeline audio docs say mp3 output re-adds priming padding at every edge and recommend wav, which is pcm_s16le and sample-exact, whenever a file will be joined again or drives lip-sync. Per-sentence slices also need wav or raw.

The defaults and the slice rule come from the Sume API reference, and the padding note and output limits from Timeline audio and Timeline 1.0, read on 2026-10-03. I did not measure the padding in milliseconds, so I do not give one.

What changes between the two?

Mostly what you can do next. File size is secondary: it is simple arithmetic from the format.

Narration formats on Sume, from the API reference and Timeline audio docs, read 2026-10-03.
Questionwavmp3
TTS defaultPass it explicitlyYes, 44,100 Hz, 128 kbps
Per-sentence slices (emit_audio)Yes, wav or rawTimings only, no slices
Join or split edgesSample-exactPriming padding added at every edge
Size of 60 seconds, one channelAbout 5,292 KB at 44.1 kHz 16-bitAbout 960 KB at 128 kbps

How do I ask for wav?

Set output_format to a wav container with pcm_s16le and 44,100 Hz, the same recipe Sume's own tool descriptions give for a voice spine. This script prints the body and the size of a take at different lengths.

RATE = 0.0475 / 1000
def wav_bytes(seconds, rate=44100, channels=1):
    return seconds * rate * 2 * channels   # pcm_s16le is 2 bytes per sample
def mp3_bytes(seconds, bit_rate=128000):
    return seconds * bit_rate // 8
for s in (30, 60, 180):
    print(s, "s: wav", wav_bytes(s) // 1000, "KB per channel; mp3", mp3_bytes(s) // 1000, "KB")

What about lip-sync and avatars?

The same rule holds there. Timeline audio's docs say to keep wav when a file drives lip-sync, and the TTS reference says to pass wav, pcm_s16le and 44,100 Hz explicitly when the audio feeds an avatar mux. So one wav take serves the join, the split and the avatar clip, and you convert once at the end if you need a smaller file.

When is mp3 still the right choice?

When the file is final and goes straight to a listener, mp3 is five to six times smaller. If you must deliver mp3 after editing, cut and join in wav, then request mp3 as the last step with Timeline audio's output.format, so the padding is added once. A join or split is one $0.01 job whatever the format, and the produced audio is capped at 1,800 seconds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume