Voiceover for shorts: one sentence per shot using TTS sentence slices
Ask TTS for timestamps.words and sentence segmentation and Sume returns gapless per-sentence slices, so each shot in a short gets its own audio file.

For a short with one idea per shot, ask TTS for sentence segmentation and you get the audio already cut. Send timestamps: {"words": true} and segmentation: {"mode": "sentence"}, and the completed job returns segments[] that are gapless: each segment's end equals the next one's start. The shot lengths come out of the audio, so the picture follows the voice and not the other way round.
The request
Segmentation requires timestamps.words to be true. Audio slices come back when emit_audio is on, and that needs a wav or raw container, so request output_format: {"container": "wav"}. The default container is mp3 at 44100 Hz and 128 kbps, which is fine for a final file but not for slices.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: shorts-slices-001" \
-d '{"transcript": "Meet the bag. It folds flat. It holds ten pounds.", "voice": {"id": "'"$VOICE_ID"'"}, "output_format": {"container": "wav"}, "timestamps": {"words": true}, "segmentation": {"mode": "sentence", "emit_audio": true}}'Cut feel
The cut falls a set number of milliseconds after the last word of each sentence, so the next segment absorbs the pause. The default boundary_lead_ms is 70, and it can range from 0 to 500. A shorter lead feels punchier and a longer one leaves room for a breath.
From segments to shots
Map each segment to one shot: the shot's duration is the segment's length. Pass the segments to a Timeline 1.0 render as video[] slots with start and duration, and the join to audio.url. If you ever need the slices as reusable files, timeline audio split handles up to 20 ranges.
What you get back, read 2026-10-06:
| Field | Value | Note |
|---|---|---|
| timestamps.words | true | Required for segmentation |
| segmentation.mode | sentence | Only mode in v1 |
| boundary_lead_ms | 0 to 500, default 70 | Lead after the last word |
| emit_audio | default true | Slices need wav or raw |
When slices are not enough
Sentence cuts follow punctuation. A script with no full stops gives you one huge segment, so write for the ear: short sentences, one idea each, and a full stop where you want a cut.
If a shot needs to be longer than its sentence, hold the picture and not the audio. A timeline slot can last longer than the line, and the voice simply ends early. Check the total against the 1200 second limit per TTS job.
Keep sentences short, because one long sentence is one long shot. If a script has more than the 1200 seconds of audio TTS allows per job, split it by chapter and run one job per chapter.
Sources
Related posts
More in Developers
- waitForJob on a failed Sume image job: read the record, skip /result
waitForJob resolves for failed and canceled jobs instead of throwing, and /result returns 409 job_not_completed for them. How to branch on job.status.
- Wan 3.0 lists 20 references; Sume caps 10 images, 5 videos, 5 audio
Wan 3.0's own page says up to 20 reference assets. Sume splits that into 10 images, 5 videos and 5 audio clips. Validate a request offline in Python first.
- What a 30-second Wan 3.0 job reserves on Sume: list times 1.25
Sume reserves the provider list price times 1.25 at submit. Work out the hold for a 30-second Wan 3.0 clip and other models in Python.
- Sume webhook beat your database write: handle an unknown job_id
A fast Sume job can send its webhook before your submit handler commits the job id. Park the event in an inbox table and reconcile it, in runnable Python.
Written by Sume