Cut a voiceover into sentence clips with TTS segmentation
Sume TTS returns gapless sentence segments, cutting 70 ms after each last word by default. Per-segment audio needs wav or raw; mp3 returns timings only.

Ask for sentence segmentation in the TTS request: set timestamps.words to true and segmentation.mode to "sentence". The result carries gapless segments, and with a wav or raw container each segment can include its own sample-exact audio_url. If you choose mp3, you still get the timings but no per-segment audio.
This is the cleanest way to get one clip per sentence for scene cutting, dubbing alignment or per-line captions without slicing the file yourself.
The fields
The TTS contract describes segmentation as optional, requiring timestamps.words: true. Mode is sentence, the only value in v1. Segments are gapless, meaning segment[i].end equals segment[i+1].start. A boundary_lead_ms value from 0 to 500, default 70, sets how many milliseconds after the last word of a sentence the cut falls; the next segment absorbs the pause. emit_audio defaults to true.
| Field | Allowed values | Default | Effect |
|---|---|---|---|
| timestamps.words | true or false | off | Word start and end times in the result |
| segmentation.mode | sentence | none | Gapless sentence segments |
| segmentation.boundary_lead_ms | 0 to 500 | 70 | Pause kept after the last word before the cut |
| segmentation.emit_audio | true or false | true | Per-segment audio_url when the container is wav or raw |
Choosing the lead
A small lead trims tight and suits fast cuts where every sentence starts a new scene. A larger lead leaves breathing room at the end of each clip so a sentence does not sound clipped when played alone. Because the pause moves into the next segment rather than being deleted, the clips still add up to the full recording.
- Keep 70 ms for rapid scene changes.
- Raise toward 200 to 300 ms for narration that will be played clip by clip.
- Use 0 only when you will add your own padding downstream.
- Test one sentence that ends on a plosive, since hard consonants show clipping first.
Pair it with the container
Set output_format.container to wav for per-segment audio, since slices from mp3 are not emitted. If you need mp3 delivery in the end, slice in wav, join as needed with timeline audio, and encode once at the final step; mp3 adds encoder padding at every edge.
The word timestamps are still there for captions. Feeding those words to the captions endpoint avoids recognition, as the timestamps post describes.
When not to segment
If you only need one finished voiceover under a video, skip segmentation; the extra output is more to store and check. Use it when sentences are units of work: one image or clip per sentence, one translation unit per sentence, or one caption block per sentence.
Related posts
More in Developers
- DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.
- Deno 2.9 Deno.test.each: a case table for a Sume webhook verifier
Deno 2.9 adds Deno.test.each. Table-test a sume-v1 verifier: valid, rotated, empty secret, stale timestamp and tampered body, with WebCrypto only.
- Deno 2.9 t.assertSnapshot: contract-test the Sume status envelope
Deno 2.9 builds assertSnapshot into the test context. Snapshot the field names and types of GET /v1/jobs/:id/status for a completed job to catch API drift.
- A dry-run flag for Sume API calls: print the request, skip the spend
Add DRY_RUN to code that calls the Sume API: build the body, key and spend cap, print them, and send nothing. Review a batch before it costs money.
Written by Sume