Sume TTS word timestamps and sentence slices: the 70 ms boundary rule
Ask Sume TTS for timestamps.words and segmentation.mode sentence to get gapless sentence segments. The cut lands 70 ms after the last word unless you change it.

Set timestamps.words to true to get word start and end times on the finished job. Add segmentation.mode of sentence and you also get gapless sentence segments, each cut 70 milliseconds after its last word by default. You can move that cut with boundary_lead_ms, from 0 to 500.
Segmentation needs word timestamps, and per-segment audio files need a wav or raw container.
The options
These come from the TTS 1.0 contract.
| Field | Value | Effect |
|---|---|---|
| timestamps.words | true or false | Adds monotonic words[] with start and end seconds |
| segmentation.mode | sentence (only mode in v1) | Adds gapless segments[]; requires timestamps.words |
| segmentation.boundary_lead_ms | 0 to 500, default 70 | Pause after the last word before the cut; the next segment absorbs the rest |
| segmentation.emit_audio | true by default | With wav or raw, each segment has a sample-exact audio_url |
| output_format.container | mp3, wav, raw | mp3 returns timings without segment audio_url |
What you can build
Sentence slices turn one read into many assets.
- Per-sentence audio for a quiz, a game or an app that plays lines separately.
- Captions and karaoke highlighting from word times.
- Chapter markers computed from the end of each segment.
- A cut list for a video editor, with each sentence as a clip start and duration.
Choosing the lead
Because segments are gapless, segment[i].end equals segment[i+1].start. The lead only decides where the silence between sentences is assigned. A larger value keeps trailing breath in the first sentence, and a smaller one hands more of the pause to the next. Start at the default of 70 ms, and listen to a few joins before changing it.
Cost is unchanged by timing options. You pay the TTS character charge for one job, rounded up to the cent, instead of one job per sentence, which matters because short jobs each hit the 1-cent minimum.
Stitch them back
If you slice a read and later want a different order, timeline audio can join up to 20 Sume-hosted parts into one gapless wav for $0.01. Keep the segments in wav so no encoder padding creeps in at the seams.
Sources
Related posts
More in Media tools
- Test a Kling motion control look on 10 seconds for $1.575, then run 30
On Sume a 10-second Kling 3.0 motion control test costs $1.575 and the full 30-second run $4.725. How to cut the test clip and what it can and cannot tell you.
- Threads video over 5 minutes: split an AI video with Sume video trim
Threads takes videos up to 5 minutes. How to cut a longer AI video into parts with POST /v1/video-trim, and what each cut costs on Sume.
- Three-minute Short: songs cap at 90 seconds, so plan the music bed
YouTube says most songs run up to 90 seconds in a 3-minute Short, some 60 or 30. Build the 180-second timeline with a voice spine and a looping, ducked bed.
- TikTok ad creatives prohibit watermarks: sample corner stills first
TikTok's reservation spec prohibits watermarks, TikTok's own included. Pull stills with video-inspect and crop the four corners before a reused clip goes up.
Written by Sume