Trim a Reel clip's dead start and end from first and last word

Take the first word's start and the last word's end from STT, then add 0.15 s lead and 0.30 s tail: a 29.9 s clip becomes 1.69-27.92 s for $0.03.

4 min readSume
All posts

The simplest cut a short-form clip needs is the dead air at both ends, before you speak and after you stop. Run STT once, take the first word's start and the last word's end, widen them by a lead (0.15 s) and a tail (0.30 s) and trim once: a 29.9 s clip with words at 1.84 s and 27.62 s becomes 1.69 to 27.92 s, a $0.03 job including STT.

The numbers

Example numbers, not a measurement. Keep full precision from the transcript; rounding the start up can clip the first syllable.

Dead-start and dead-end trim for one clip (hypothetical clip, arithmetic only)
QuantityValue
Clip length29.90 s
First word starts1.84 s
Last word ends27.62 s
Trim start (1.84 - 0.15)1.69 s
Trim end (27.62 + 0.30)27.92 s
Kept26.23 s
Removed3.67 s
Cost: STT $0.01 + trim $0.02$0.03

Code

The tail is longer than the lead because words trail off; the trim end is clamped to the clip by the API, which adds the warning trim_clamped_to_source if you overshoot.

def dead_ends(words, lead=0.15, tail=0.30, clip_s=None):
    timed = [w for w in words if "start" in w and "end" in w]
    if not timed:
        return None
    start = max(timed[0]["start"] - lead, 0)
    end = timed[-1]["end"] + tail
    if clip_s:
        end = min(end, clip_s)
    return {"start": round(start, 3), "end": round(end, 3)}

Gotchas

  • Silent clips have no words; STT on a silent clip fails with inspect_source_has_no_audio, so read probe.has_audio first.
  • A speech-to-text pass reserves one minute when you omit duration_seconds, so a 30 s clip costs $0.01.
  • The result must be at least 0.2 s, which a spoken clip always clears.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume