TTS boundary_lead_ms: why sentence slices end 70 ms after a word
Sume TTS sentence segments cut boundary_lead_ms after a sentence's last word: 70 ms by default, 0 to 500. The next segment absorbs the pause, and no gap opens.

segmentation.boundary_lead_ms sets how many milliseconds Sume TTS keeps after the last word of a sentence before it cuts. The default is 70, the allowed range is 0 to 500, and the next sentence's segment takes the pause that follows. That is why segment ends sit slightly after the spoken word, and why every segments[i].end equals segments[i+1].start.
What does the field do, exactly?
The OpenAPI description reads: milliseconds after the last word of a sentence before the cut; the next segment absorbs the pause; default 70. It is an integer with minimum 0 and maximum 500, and it sits beside mode: "sentence" in the segmentation object, which needs timestamps.words: true.
The 70 ms default is the lead the docs call the post-word boundary rule. A value outside 0 to 500 is refused with the message that segmentation.boundary_lead_ms must be an integer between 0 and 500.
| Value | Meaning | Result |
|---|---|---|
| omitted | Default of 70 ms after the last word | Cut sits 70 ms past the word |
0 | Cut right after the last word's timing | The pause goes to the next segment |
500 | Maximum, half a second of lead | More trailing room in each slice |
501 or 1.5 | Outside the schema | Rejected with a validation error |
Why would I change it?
Raise it when a slice clips the tail of a word. Lower it when you join slices to a visual cut and want less trailing air in each clip. Because the segments are gapless, whatever the lead adds to one slice comes off the pause at the start of the next, so the total length does not change.
Cartesia's changelog (read 2026-10-02) says timestamps behave on Sonic 3.6 as they do on Sonic 3.5. Word timings feed this rule, so a model switch is not a reason to retune it.
Do I get audio files for the slices?
Only with wav or raw output. With the default mp3 you get timings and no per-segment audio_url; see the emit_audio rule. When slices are made they are sample-exact.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lead-150-001" \
-d '{
"transcript": "One. Two. Three.",
"avatar_handle": "@your_avatar",
"language": "en",
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence", "boundary_lead_ms": 150 }
}'Sources
Related posts
More in Developers
- tts_duration_exceeded: split long scripts for Sume TTS
Sume TTS fails with tts_duration_exceeded when audio would pass 1,200 seconds, and takes at most 20,000 characters. Split rules and how to join the parts.
- TTS language omitted: Spanish text comes out with an English default
On Sume TTS, an omitted language defaults to English at the provider, with a ko or ja fallback only for Hangul or kana text. Set language for all others.
- TTS sentence slices have no audio_url: emit_audio needs wav or raw
With segmentation on, Sume TTS only returns a sample-exact audio_url per sentence for wav or raw output. For mp3 you get timings and a warning instead.
- TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
Sume TTS marks the slow, normal and fast speed enum deprecated. Send generation_config.speed from 0.6 to 1.5, plus volume 0.5 to 2 and an emotion string.
Written by Sume