Trim the silent lead-in: first spoken word from STT, then video trim
A clip that starts with two seconds of nothing loses viewers. Read the first word's start time from STT, then cut with video trim, exact or keyframe.

How do you remove the silence before someone starts talking? Ask speech-to-text when the first word begins, then cut the video from just before that time. Sume's STT returns words[] with start and end seconds, and POST /v1/video-trim cuts [start, end) into a new MP4.
Fast first partials are a selling point for live models: Microsoft says MAI-Transcribe-2-Streaming produces first text just over 100 ms after audio arrives (source, read 2026-10-04). For a finished clip you do not need that, only an accurate first timestamp.
Get the first word
Video inspect with transcribe: true runs Sume STT on the clip's audio and returns transcript with text and words[]. The transcript adds the STT rate, documented as $0.01 per audio minute, to the inspect reservation. Word items can include a type of word or spacing when the provider supplies one, so skip any spacing token when you look for the first word.
Cut with a small pad
Start the cut a little before the word so the breath survives. Use start plus exactly one of end or duration; sending both returns video_trim_range_conflict. The default precision: exact is a frame-accurate re-encode. precision: keyframe is a stream copy whose cut may begin a GOP early, so re-base against actual_start_seconds in the result.
A trim costs $0.02 per job per the video trim guide; confirm live in GET /v1/catalog.
import os, requests
first_word_start = 2.31 # words[0].start from the STT result
clip_duration = 24.0 # probe.duration_seconds from video inspect
pad = 0.15
start = max(0.0, first_word_start - pad)
r = requests.post(
"https://api.sume.com/v1/video-trim",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "trim-leadin-001"},
json={"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"start": round(start, 2), "end": clip_duration,
"precision": "exact"})
print(r.status_code)
Edge cases
If end runs past the source it clamps and the result warns trim_clamped_to_source. If the clip has no speech at all, the transcript has no first word; handle that case instead of cutting at zero. Output must be at least 0.2 seconds and at most 900 seconds.
Sources
Related posts
More in Media tools
- Video analyses: include_transcript false leaves scene audio null
On the legacy Sume video-analyses resource, scene audio is null unless include_transcript was true. Handle both shapes, and know what replaces it for new work.
- Video analyses keyframes: one still per second per scene
A legacy Sume video analysis gives a still for each whole second of a scene plus a keyframe_url near 40%. How stills fail and how to pick your own cover.
- Video filter /check: submit or fix and recheck
POST /v1/video-filter/check is unbilled and returns valid, diagnostics and a next_action of submit_video_filter or fix_program_and_recheck. Branch on it.
- Video inspect fast seek: requested_times vs sample_times
A fast-seek inspect grid returns requested_times and sample_times. Use sample_times for what the tiles show, then trim from them, never from the request.
Written by Sume