Trim the silent lead-in: first spoken word from STT, then video trim

A clip that starts with two seconds of nothing loses viewers. Read the first word's start time from STT, then cut with video trim, exact or keyframe.

5 min readSume
All posts

How do you remove the silence before someone starts talking? Ask speech-to-text when the first word begins, then cut the video from just before that time. Sume's STT returns words[] with start and end seconds, and POST /v1/video-trim cuts [start, end) into a new MP4.

Fast first partials are a selling point for live models: Microsoft says MAI-Transcribe-2-Streaming produces first text just over 100 ms after audio arrives (source, read 2026-10-04). For a finished clip you do not need that, only an accurate first timestamp.

Get the first word

Video inspect with transcribe: true runs Sume STT on the clip's audio and returns transcript with text and words[]. The transcript adds the STT rate, documented as $0.01 per audio minute, to the inspect reservation. Word items can include a type of word or spacing when the provider supplies one, so skip any spacing token when you look for the first word.

Cut with a small pad

Start the cut a little before the word so the breath survives. Use start plus exactly one of end or duration; sending both returns video_trim_range_conflict. The default precision: exact is a frame-accurate re-encode. precision: keyframe is a stream copy whose cut may begin a GOP early, so re-base against actual_start_seconds in the result.

A trim costs $0.02 per job per the video trim guide; confirm live in GET /v1/catalog.

import os, requests

first_word_start = 2.31      # words[0].start from the STT result
clip_duration = 24.0         # probe.duration_seconds from video inspect
pad = 0.15
start = max(0.0, first_word_start - pad)
r = requests.post(
    "https://api.sume.com/v1/video-trim",
    headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
             "Idempotency-Key": "trim-leadin-001"},
    json={"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
          "start": round(start, 2), "end": clip_duration,
          "precision": "exact"})
print(r.status_code)

Edge cases

If end runs past the source it clamps and the result warns trim_clamped_to_source. If the clip has no speech at all, the transcript has no first word; handle that case instead of cutting at zero. Output must be at least 0.2 seconds and at most 900 seconds.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume