Voiceover for shorts: one sentence per shot using TTS sentence slices

Ask TTS for timestamps.words and sentence segmentation and Sume returns gapless per-sentence slices, so each shot in a short gets its own audio file.

4 min readSume
All posts

For a short with one idea per shot, ask TTS for sentence segmentation and you get the audio already cut. Send timestamps: {"words": true} and segmentation: {"mode": "sentence"}, and the completed job returns segments[] that are gapless: each segment's end equals the next one's start. The shot lengths come out of the audio, so the picture follows the voice and not the other way round.

The request

Segmentation requires timestamps.words to be true. Audio slices come back when emit_audio is on, and that needs a wav or raw container, so request output_format: {"container": "wav"}. The default container is mp3 at 44100 Hz and 128 kbps, which is fine for a final file but not for slices.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: shorts-slices-001" \
  -d '{"transcript": "Meet the bag. It folds flat. It holds ten pounds.", "voice": {"id": "'"$VOICE_ID"'"}, "output_format": {"container": "wav"}, "timestamps": {"words": true}, "segmentation": {"mode": "sentence", "emit_audio": true}}'

Cut feel

The cut falls a set number of milliseconds after the last word of each sentence, so the next segment absorbs the pause. The default boundary_lead_ms is 70, and it can range from 0 to 500. A shorter lead feels punchier and a longer one leaves room for a breath.

From segments to shots

Map each segment to one shot: the shot's duration is the segment's length. Pass the segments to a Timeline 1.0 render as video[] slots with start and duration, and the join to audio.url. If you ever need the slices as reusable files, timeline audio split handles up to 20 ranges.

What you get back, read 2026-10-06:

TTS sentence slices, read 2026-10-06
FieldValueNote
timestamps.wordstrueRequired for segmentation
segmentation.modesentenceOnly mode in v1
boundary_lead_ms0 to 500, default 70Lead after the last word
emit_audiodefault trueSlices need wav or raw

When slices are not enough

Sentence cuts follow punctuation. A script with no full stops gives you one huge segment, so write for the ear: short sentences, one idea each, and a full stop where you want a cut.

If a shot needs to be longer than its sentence, hold the picture and not the audio. A timeline slot can last longer than the line, and the voice simply ends early. Check the total against the 1200 second limit per TTS job.

Keep sentences short, because one long sentence is one long shot. If a script has more than the 1200 seconds of audio TTS allows per job, split it by chapter and run one job per chapter.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume