Sume TTS sentence segments: boundary_lead_ms 0 vs 500, 12 lines
Ask Sume TTS for words and sentence segmentation and you get gapless wav slices. boundary_lead_ms (default 70) moves each cut 0 to 500 ms after the last word.

To get one audio file per sentence from a single Sume TTS job, send timestamps.words: true and segmentation: {"mode": "sentence"} with a wav output. The result carries gapless segments[] (each segment's end equals the next one's start), and boundary_lead_ms sets how many milliseconds past a sentence's last word the cut falls. The default is 70 and the allowed range is 0 to 500.
These rules are from the TTS request schema in the Sume repository (read 2026-10-09). A 12-line ad read costs the same as the unsegmented read: $0.0475 per 1,000 characters.
What the fields do
segmentation requires timestamps.words: true. emit_audio defaults to true and gives each segment a sample-exact audio_url, but only when output_format.container is wav or raw; with mp3 you get the timings and no slice URLs. The next segment absorbs the pause, so nothing is lost between slices.
| Field | Values | Effect |
|---|---|---|
| timestamps.words | true | Returns words[] with start and end seconds |
| segmentation.mode | sentence | Required; only value in v1 |
| segmentation.boundary_lead_ms | 0 to 500, default 70 | Milliseconds after a sentence's last word before the cut |
| segmentation.emit_audio | true by default | Sample-exact audio_url per segment, wav or raw only |
| output_format.container | mp3, wav, raw | Use wav for slices |
Choosing 0 or 500
At 0 the cut lands on the last word's end, which gives the tightest clip and the highest risk of clipping a trailing consonant. At 500 each slice keeps half a second of tail, which is safer for a clip you will play alone but adds to the pause when you join slices back. The default of 70 is a compromise. For a 12-line ad read that you will cut to picture, start at the default and raise it only if you hear a clipped ending.
Because the next segment absorbs the pause, raising the lead also shortens the start of the following clip by the same amount, and the total stays the length of the take.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ad-read-12-lines-001" \
-d '{
"transcript": "Line one. Line two. Line three.",
"avatar_handle": "studio_presenter",
"output_format": {"container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le"},
"timestamps": {"words": true},
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70, "emit_audio": true}
}'Using the slices
Each wav slice is a Sume media file, so you can place slices in a render with audio.parts[] in Timeline 1.0 or join some of them with timeline audio for $0.01 a job. Use the segment offsets when you re-base video starts.
Checking the slices
Before you publish, listen to the first and last slice and one in the middle, and check that the sum of segment durations equals the take length. Because segments are gapless, that sum should match to the sample.
If a slice sounds clipped, raise boundary_lead_ms in steps of 50 and re-run only that script. Since the price depends on the characters and not on the segmentation options, a re-run of a 1,200-character read costs 1.2 x $0.0475 = $0.057, so it is cheap to tune.
One more constraint: segmentation returns sentence slices, so a line without terminal punctuation may merge with the next. Write each of your 12 ad lines as a full sentence ending in a period, and the slice count will match the line count.
Sources
Related posts
More in Developers
- TTS speed 1.2 turns a 30-second read into about 25 seconds
Sume TTS accepts generation_config.speed from 0.6 to 1.5. Speed changes length, not the bill: 450 characters cost $0.021375 at any speed. Test the real length.
- TTS word timings straight into captions: a 45-second ad for 33 cents
Ask Sume TTS for timestamps.words and send them as words on the caption job: no speech-to-text. 700 characters, render and captions come to $0.33.
- 20 Nano Banana 2.1 2K images: $3.00 and one jobs_result read
Twenty Nano Banana 2.1 images at 2K cost $3.00 on Sume. Over hosted MCP, one jobs_result call with 20 job_ids reads them back; failed ids are named.
- 20 Omni Flash 1.1 clips in one jobs_wait wave: $12.50 at 720p
Twenty 5-second Gemini Omni Flash 1.1 clips at 720p cost $12.50 on Sume ($0.625 each) and fit one jobs_wait call of 20 ids. Cost table and the call body.
Written by Sume