video_inspect silence_split_seconds: sentence segments for captions

How silence_split_seconds (0.2 to 3) shapes Sume video-inspect sentence segments, the 0.5 s default in the repo, and turning segments into caption cues.

5 min readSume
All posts

silence_split_seconds is an optional number from 0.2 to 3 on segmentation in a Sume video inspect request. It sets how long a pause between words must be before a sentence segment is cut where there was no punctuation to cut it. Leave it out and the worker uses 0.5 seconds, a default set in the Sume repository's STT segmentation code.

It only does anything when transcribe: true and segmentation.mode is "sentence". Sending segmentation without transcribe: true returns 400 video_inspect_transcribe_required. The segments come back alongside the words, each with an index, text, start, end and duration_seconds, and they are gapless and shaped like caption lines.

What does the silence setting actually split?

Reading the segmentation code in the repo, sentences are first closed by sentence-ending punctuation in the transcript. The silence rule is then applied to the run of words left after the last punctuated sentence end. In a transcript that carries punctuation throughout, the setting rarely changes anything. In unpunctuated narration, where the speech-to-text returns few or no full stops, the whole run is split at every gap of at least the threshold.

The code's own comment explains the 0.5 second default: long enough to ignore ordinary within-sentence hesitation, short enough to break up unpunctuated live-commerce narration. That is a design note about the default, not a promise about every language or speaker. The public docs give only the range, so treat the behaviour above as a repo-level detail that could change.

Which value should you try?

The suggested ranges are starting points we derived from how the rule works, not documented recommendations. Always check segments on a representative clip before you batch.

silence_split_seconds symptoms and adjustments for Sume video inspect, read 2026-10-02
What you seeTryWhy
One very long segment on unpunctuated speechLower toward 0.3 to 0.5More pauses count as boundaries
Lines chopped mid-thought on a slow speakerRaise toward 0.8 to 1.5Short hesitations stop splitting
Punctuated transcript, setting seems ignoredLeave the defaultPunctuation closes sentences before the silence rule runs
Need the longest possible linesUse 3 and cap line length in the caption callOnly a 3 second gap splits unpunctuated runs

How do segments become burned-in captions?

Sume's caption job accepts cues (or segments) with text, start and end in seconds and burns exactly that copy with no second speech-to-text pass. The inspect segments already have those three fields, so the mapping is a list comprehension. This script transcribes with a custom split, reads the resource and prints cues ready to send to POST /v1/video-captions. It uses the standard library only.

import json, os, urllib.request

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
     "Content-Type": "application/json"}

def call(method, url, body=None, key=None):
    h = dict(H)
    if key:
        h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    req = urllib.request.Request(url, data, h, method=method)
    return json.load(urllib.request.urlopen(req))

clip = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
r = call("POST", "https://api.sume.com/v1/video-inspect", {
    "video_url": clip, "frames": False, "transcribe": True,
    "segmentation": {"mode": "sentence", "silence_split_seconds": 0.8},
}, key="inspect-seg-001")
res = call("GET", "https://api.sume.com/v1/video-inspect/" + r["request_id"])
res = res.get("video_inspect", res)
cues = [{"text": s["text"], "start": s["start"], "end": s["end"]}
        for s in res["transcript"]["segments"]]
print(json.dumps(cues[:3], indent=1))

What does it cost and where does it fail?

The probe is unbilled and the transcript reserves $0.01 per audio minute (confirm in GET /v1/catalog). If you omit duration_seconds, one minute is reserved; the maximum hint is 600 seconds. A clip with no audio track fails with inspect_source_has_no_audio, so check probe.has_audio first with frames: false. The default mode is sync with a 30 second wait: you get 200 with the finished inspect or 202 and a job to poll, which is why the script reads the resource afterwards.

The clip must already be on media.sume.com for your workspace; import it with POST /v1/media-imports first. The later caption call is different and takes a public HTTPS URL, which the video captions docs show with a media.sume.com artifact URL.

Should you cut cues again before burning them?

Often yes. Sentence segments can be long, and a caption card that holds a whole sentence is hard to read on a phone. The caption call has its own phrasing controls under design.phrasing (max_words, max_chars, pause_seconds). The docs say authored cues skip speech-to-text and are burned as authored copy, so split long segments yourself before sending them rather than relying on those controls. A simple rule is to break any segment longer than about two lines of text at the nearest word boundary and divide the time in proportion to character count. That is an approximation; for exact word timings, pass words instead of cues.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume