TTS sentence slices: segmentation, wav output and boundary_lead_ms 70

Sume TTS can return gapless sentence segments. Needs timestamps.words and wav or raw for per-sentence audio_url. A 900-character script costs $0.04275.

4 min readSume
All posts

Add segmentation {"mode": "sentence"} to a Sume TTS request, with timestamps.words true, and the result returns gapless segments[]. Use a wav or raw container to get a sample-exact audio_url per sentence; a 900-character script is $0.04275.

How the options combine

Rules are from the OpenAPI schema for segmentation.

Sume TTS segmentation (read 2026-10-09)
SettingValueEffect
modesentence (only option)Cuts at sentence ends
boundary_lead_ms0 to 500, default 70Cut falls this long after the last word; the next segment absorbs the pause
emit_audiodefault trueWith wav or raw: per-segment audio_url. With mp3: timings only

Request

Per-sentence slices are handy when each sentence maps to one scene.

{
  "transcript": "Rain starts at noon. Bring a jacket. Dry by evening.",
  "avatar_handle": "@narrator",
  "language": "en",
  "output_format": {"container": "wav", "sample_rate": 24000, "encoding": "pcm_s16le"},
  "timestamps": {"words": true},
  "segmentation": {"mode": "sentence", "boundary_lead_ms": 70}
}

Gotchas

Segmentation without timestamps.words is invalid. With mp3 you still get segment timings but no slices. Segments are gapless: segment[i].end equals segment[i+1].start, so pauses live at the start of the next sentence.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume