silence_split_seconds: tune caption line breaks from STT segments
Sume STT sentence segmentation can split on silence. silence_split_seconds takes 0.2 to 3 and returns gapless segments you can use as caption lines.

On Sume, sentence segmentation returns segments[] with no gaps, in the shape of caption lines, and silence_split_seconds sets how long a pause must be before it splits a line. The video inspect page gives the range as 0.2 to 3. A short value breaks lines at small pauses and gives more, shorter cards. A long value gives fewer, longer ones.
Where to set it
The setting is part of segmentation on a transcribing request. On video inspect you send transcribe: true, then segmentation.mode: "sentence". If you send language_code, segmentation or duration_seconds without transcribe: true, the API returns 400 video_inspect_transcribe_required.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"transcribe": true,
"segmentation": { "mode": "sentence", "silence_split_seconds": 0.6 }
}Choosing a value
The table says what the setting does in direction only. The docs do not give a recommended value, so test one clip and read the segments.
| Value | Effect on caption lines | Try it for |
|---|---|---|
| 0.2 | Splits at the smallest pauses; many short lines | Fast, clipped speech |
| 0.6 to 1 | Middle ground | Talking-head clips |
| 3 | Splits only on long pauses; long lines | Slow narration |
From segments to captions
Segments are time ranges, not sliced audio files. To burn them, map each segment to a cues entry with text, start and end, and send those to POST /v1/video-captions. Sume then burns your text at those times without running speech-to-text again. Send only one of script_text, words, cues and segments.
Transcribe adds the STT rate of $0.01 per audio minute to the inspect reservation. Without duration_seconds, Sume reserves 1 minute, and the maximum hint is 600 seconds.
Sources
Related posts
More in Developers
- Captions fail on a silent Short with caption_no_speech: send cues
A silent clip has no speech to transcribe, so Sume's captions API returns caption_no_speech. Send cues with text, start and end to burn overlay text instead.
- 16 AI shots in one Sume Timeline render: the 8-fade cap
A 16-shot cut fits one Timeline render, but fades are capped at 8 in a row and renders chunk past 12 slots. Plan the cuts, with the doc limits.
- Size a batch from generation_limits so no clip hits queue_full
Read accepted_generation_jobs_limit from a Sume submit response and slice your clips. A 50-clip batch leaves 2 for a later wave on Startup and 26 on Pro.
- Voice model updated in place with no API change: how to detect it
Nova 2 Sonic was refreshed in place in May with no API change. If a vendor can change your voice silently, log the model id and a canary clip. Sume code inside.
Written by Sume