Cut a quote clip from a long video: transcript, then video_trim

Find the sentence in a Sume transcript, then cut it with video_trim. A Python script, with timing padding, $0.02 per trim, and the 900-second output cap.

5 min readSume
All posts

Ask video_inspect for a sentence-segmented transcript, find the sentence that holds your quote, and pass its start and end to video_trim. At public rates an 8-minute talk costs about $0.08 for the transcript plus $0.02 for the cut, before compute. The script below does the lookup and the cut in under 30 lines.

Step 1: transcript with sentence segments

Send transcribe: true with segmentation.mode set to sentence and a duration_seconds hint (maximum 600). The result carries a transcript with words and sentence segments, each with start, end and text. The clip must already be a media.sume.com file in your workspace, since the API does not fetch open-internet URLs.

Step 2: the cut

video_trim takes video_url, start, and one of end or duration. The default precision is exact, a frame-accurate re-encode; keyframe is a stream copy that can begin up to one GOP early. Pad the sentence by about 0.3 seconds each side so the clip does not clip a breath. The cut runs async, so poll the job and read the result's video_url.

import os, requests
B = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def main():
    src = "https://media.sume.com/artifacts/artf_demo/talk.mp4"
    r = requests.post(B + "/v1/video-inspect", headers={**H, "Idempotency-Key": "quote-inspect-1"},
        json={"video_url": src, "frames": False, "transcribe": True, "duration_seconds": 480,
              "segmentation": {"mode": "sentence"}})
    r.raise_for_status()
    segs = r.json()["video_inspect"]["transcript"]["segments"]
    hit = next(s for s in segs if "our retention doubled" in s["text"].lower())
    start = max(0, hit["start"] - 0.3)
    end = hit["end"] + 0.3
    t = requests.post(B + "/v1/video-trim", headers={**H, "Idempotency-Key": "quote-trim-1"},
        json={"video_url": src, "start": start, "end": end})
    t.raise_for_status()
    print(t.json()["request_id"], round(end - start, 2))

if __name__ == "__main__":
    main()

Limits and costs

Check the numbers before you batch this. The source must be 1,800 seconds or less and the output 900 seconds or less. If the clip is a sync inspect that takes longer than 30 seconds, you get a 202 and must poll the job; the script above assumes a fast answer, so add polling for longer files.

Quote-clip steps, from Sume docs read 2026-10-05
StepToolPublic rateLimit
Transcriptvideo_inspect, transcribe true$0.01 per audio minuteHint up to 600 s
Cutvideo_trim$0.02 per jobOutput 0.2 to 900 s
Place in an edittimeline_createPer timeline docsUse source_in 0
Sourcemedia.sume.com clipImport first1,800 s max

Choosing padding and precision

Padding of 0.2 to 0.5 seconds suits most speech. Keep it smaller when the next sentence begins right away, or the clip will carry the first syllable of the following line. Use the default exact precision for anything you will publish, because a keyframe cut can start up to one GOP early, roughly 0 to 5 seconds on typical sources, and the result then reports a different actual_start_seconds from the one you asked for. Keyframe cuts are fine for a rough review copy, since they skip the re-encode.

When the quote is not one sentence

If the speaker splits the line across two sentences, take the start of the first and the end of the second. If the phrase is a close match rather than exact, search on a distinctive three-word run and print the hit for a human to check. Words with timings are in the transcript too, and a quote that crosses a pause is a good reason to use them, so you can cut on a word edge when a sentence boundary is wrong. Always open the finished clip before you publish a quote, since a transcript can mishear a name or a number even when the timing is right.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume