video_inspect silence_split_seconds: sentence segments for captions
How silence_split_seconds (0.2 to 3) shapes Sume video-inspect sentence segments, the 0.5 s default in the repo, and turning segments into caption cues.

silence_split_seconds is an optional number from 0.2 to 3 on segmentation in a Sume video inspect request. It sets how long a pause between words must be before a sentence segment is cut where there was no punctuation to cut it. Leave it out and the worker uses 0.5 seconds, a default set in the Sume repository's STT segmentation code.
It only does anything when transcribe: true and segmentation.mode is "sentence". Sending segmentation without transcribe: true returns 400 video_inspect_transcribe_required. The segments come back alongside the words, each with an index, text, start, end and duration_seconds, and they are gapless and shaped like caption lines.
What does the silence setting actually split?
Reading the segmentation code in the repo, sentences are first closed by sentence-ending punctuation in the transcript. The silence rule is then applied to the run of words left after the last punctuated sentence end. In a transcript that carries punctuation throughout, the setting rarely changes anything. In unpunctuated narration, where the speech-to-text returns few or no full stops, the whole run is split at every gap of at least the threshold.
The code's own comment explains the 0.5 second default: long enough to ignore ordinary within-sentence hesitation, short enough to break up unpunctuated live-commerce narration. That is a design note about the default, not a promise about every language or speaker. The public docs give only the range, so treat the behaviour above as a repo-level detail that could change.
Which value should you try?
The suggested ranges are starting points we derived from how the rule works, not documented recommendations. Always check segments on a representative clip before you batch.
| What you see | Try | Why |
|---|---|---|
| One very long segment on unpunctuated speech | Lower toward 0.3 to 0.5 | More pauses count as boundaries |
| Lines chopped mid-thought on a slow speaker | Raise toward 0.8 to 1.5 | Short hesitations stop splitting |
| Punctuated transcript, setting seems ignored | Leave the default | Punctuation closes sentences before the silence rule runs |
| Need the longest possible lines | Use 3 and cap line length in the caption call | Only a 3 second gap splits unpunctuated runs |
How do segments become burned-in captions?
Sume's caption job accepts cues (or segments) with text, start and end in seconds and burns exactly that copy with no second speech-to-text pass. The inspect segments already have those three fields, so the mapping is a list comprehension. This script transcribes with a custom split, reads the resource and prints cues ready to send to POST /v1/video-captions. It uses the standard library only.
import json, os, urllib.request
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
def call(method, url, body=None, key=None):
h = dict(H)
if key:
h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(url, data, h, method=method)
return json.load(urllib.request.urlopen(req))
clip = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
r = call("POST", "https://api.sume.com/v1/video-inspect", {
"video_url": clip, "frames": False, "transcribe": True,
"segmentation": {"mode": "sentence", "silence_split_seconds": 0.8},
}, key="inspect-seg-001")
res = call("GET", "https://api.sume.com/v1/video-inspect/" + r["request_id"])
res = res.get("video_inspect", res)
cues = [{"text": s["text"], "start": s["start"], "end": s["end"]}
for s in res["transcript"]["segments"]]
print(json.dumps(cues[:3], indent=1))What does it cost and where does it fail?
The probe is unbilled and the transcript reserves $0.01 per audio minute (confirm in GET /v1/catalog). If you omit duration_seconds, one minute is reserved; the maximum hint is 600 seconds. A clip with no audio track fails with inspect_source_has_no_audio, so check probe.has_audio first with frames: false. The default mode is sync with a 30 second wait: you get 200 with the finished inspect or 202 and a job to poll, which is why the script reads the resource afterwards.
The clip must already be on media.sume.com for your workspace; import it with POST /v1/media-imports first. The later caption call is different and takes a public HTTPS URL, which the video captions docs show with a media.sume.com artifact URL.
Should you cut cues again before burning them?
Often yes. Sentence segments can be long, and a caption card that holds a whole sentence is hard to read on a phone. The caption call has its own phrasing controls under design.phrasing (max_words, max_chars, pause_seconds). The docs say authored cues skip speech-to-text and are burned as authored copy, so split long segments yourself before sending them rather than relying on those controls. A simple rule is to break any segment longer than about two lines of text at the nearest word boundary and divide the time in proportion to character count. That is an approximation; for exact word timings, pass words instead of cues.
Sources
Related posts
More in Developers
- Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.
- Which Sume audio endpoint to call: TTS, STT, music, detach, timeline
A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.
- Sume video tools: public URL or media import first? Per tool
Video captions takes a public HTTPS URL; trim, filter, inspect, frames, compose and detach need a workspace media.sume.com clip. A tool-by-tool input guide.
- Which voice does my avatar speak with? Check voice.status is ready
Sume TTS speaks in an avatar's voice when voice.status is ready. List avatars, check voice.status, then send avatar_id or avatar_handle on the TTS request.
Written by Sume