Sume TTS sentence segments: boundary_lead_ms 0 vs 500, 12 lines

Ask Sume TTS for words and sentence segmentation and you get gapless wav slices. boundary_lead_ms (default 70) moves each cut 0 to 500 ms after the last word.

5 min readSume
All posts

To get one audio file per sentence from a single Sume TTS job, send timestamps.words: true and segmentation: {"mode": "sentence"} with a wav output. The result carries gapless segments[] (each segment's end equals the next one's start), and boundary_lead_ms sets how many milliseconds past a sentence's last word the cut falls. The default is 70 and the allowed range is 0 to 500.

These rules are from the TTS request schema in the Sume repository (read 2026-10-09). A 12-line ad read costs the same as the unsegmented read: $0.0475 per 1,000 characters.

What the fields do

segmentation requires timestamps.words: true. emit_audio defaults to true and gives each segment a sample-exact audio_url, but only when output_format.container is wav or raw; with mp3 you get the timings and no slice URLs. The next segment absorbs the pause, so nothing is lost between slices.

Segmentation fields on Sume TTS, from the request schema as of 2026-10-09
FieldValuesEffect
timestamps.wordstrueReturns words[] with start and end seconds
segmentation.modesentenceRequired; only value in v1
segmentation.boundary_lead_ms0 to 500, default 70Milliseconds after a sentence's last word before the cut
segmentation.emit_audiotrue by defaultSample-exact audio_url per segment, wav or raw only
output_format.containermp3, wav, rawUse wav for slices

Choosing 0 or 500

At 0 the cut lands on the last word's end, which gives the tightest clip and the highest risk of clipping a trailing consonant. At 500 each slice keeps half a second of tail, which is safer for a clip you will play alone but adds to the pause when you join slices back. The default of 70 is a compromise. For a 12-line ad read that you will cut to picture, start at the default and raise it only if you hear a clipped ending.

Because the next segment absorbs the pause, raising the lead also shortens the start of the following clip by the same amount, and the total stays the length of the take.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ad-read-12-lines-001" \
  -d '{
    "transcript": "Line one. Line two. Line three.",
    "avatar_handle": "studio_presenter",
    "output_format": {"container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le"},
    "timestamps": {"words": true},
    "segmentation": {"mode": "sentence", "boundary_lead_ms": 70, "emit_audio": true}
  }'

Using the slices

Each wav slice is a Sume media file, so you can place slices in a render with audio.parts[] in Timeline 1.0 or join some of them with timeline audio for $0.01 a job. Use the segment offsets when you re-base video starts.

Checking the slices

Before you publish, listen to the first and last slice and one in the middle, and check that the sum of segment durations equals the take length. Because segments are gapless, that sum should match to the sample.

If a slice sounds clipped, raise boundary_lead_ms in steps of 50 and re-run only that script. Since the price depends on the characters and not on the segmentation options, a re-run of a 1,200-character read costs 1.2 x $0.0475 = $0.057, so it is cheap to tune.

One more constraint: segmentation returns sentence slices, so a line without terminal punctuation may merge with the next. Write each of your 12 ad lines as a full sentence ending in a period, and the slice count will match the line count.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume