Cut a voiceover into sentence clips with TTS segmentation

Sume TTS returns gapless sentence segments, cutting 70 ms after each last word by default. Per-segment audio needs wav or raw; mp3 returns timings only.

4 min readSume
All posts

Ask for sentence segmentation in the TTS request: set timestamps.words to true and segmentation.mode to "sentence". The result carries gapless segments, and with a wav or raw container each segment can include its own sample-exact audio_url. If you choose mp3, you still get the timings but no per-segment audio.

This is the cleanest way to get one clip per sentence for scene cutting, dubbing alignment or per-line captions without slicing the file yourself.

The fields

The TTS contract describes segmentation as optional, requiring timestamps.words: true. Mode is sentence, the only value in v1. Segments are gapless, meaning segment[i].end equals segment[i+1].start. A boundary_lead_ms value from 0 to 500, default 70, sets how many milliseconds after the last word of a sentence the cut falls; the next segment absorbs the pause. emit_audio defaults to true.

TTS segmentation options from the Sume OpenAPI contract (read 2026-10-03)
FieldAllowed valuesDefaultEffect
timestamps.wordstrue or falseoffWord start and end times in the result
segmentation.modesentencenoneGapless sentence segments
segmentation.boundary_lead_ms0 to 50070Pause kept after the last word before the cut
segmentation.emit_audiotrue or falsetruePer-segment audio_url when the container is wav or raw

Choosing the lead

A small lead trims tight and suits fast cuts where every sentence starts a new scene. A larger lead leaves breathing room at the end of each clip so a sentence does not sound clipped when played alone. Because the pause moves into the next segment rather than being deleted, the clips still add up to the full recording.

  • Keep 70 ms for rapid scene changes.
  • Raise toward 200 to 300 ms for narration that will be played clip by clip.
  • Use 0 only when you will add your own padding downstream.
  • Test one sentence that ends on a plosive, since hard consonants show clipping first.

Pair it with the container

Set output_format.container to wav for per-segment audio, since slices from mp3 are not emitted. If you need mp3 delivery in the end, slice in wav, join as needed with timeline audio, and encode once at the final step; mp3 adds encoder padding at every edge.

The word timestamps are still there for captions. Feeding those words to the captions endpoint avoids recognition, as the timestamps post describes.

When not to segment

If you only need one finished voiceover under a video, skip segmentation; the extra output is more to store and check. Use it when sentences are units of work: one image or clip per sentence, one translation unit per sentence, or one caption block per sentence.

Related posts

More in Developers

All Developers posts

Written by Sume