Speedrun PB video: trim the run, burn split times as cues

Cut a personal-best run out of a long capture with video trim, then burn split times as timed caption cues: two jobs, $0.22 at the listed rates.

4 min readSume
All posts

To make a speedrun personal-best video, import your capture to Sume, cut the run with video trim (a start plus a duration, $0.02), then burn the split times with video captions using authored cues, each a piece of text with a start and an end ($0.20 for a clip up to 60 seconds). Total at the listed rates: $0.22. The splits are text you already know, so no speech-to-text runs and a silent game capture is not a problem.

Why author the splits instead of transcribing them?

The standard caption path listens for speech. The captions page says speech captions work only when the clip has audible speech, and a silent clip fails with caption_no_speech and the next action use_overlay_captions. A speedrun is the typical case: game audio, maybe a muted microphone, and nothing to transcribe. Sending cues (or segments) skips that step and burns exactly the text you wrote at exactly the times you gave.

Authored text also means the numbers are right. You type 1:18.9 once, and it renders as 1:18.9, instead of depending on what a transcription heard.

What do the two jobs cost?

The trim is plain worker work with no model inference. The caption job is a flat fixed estimate.

Per-job rates from docs.sume.com, read 2026-10-11
StepRouteRateWhat you send
Cut the runPOST /v1/video-trim$0.02 per jobvideo_url, start, duration
Burn the splitsPOST /v1/video-captions$0.20 per job, clips up to 60 svideo_url, style, cues
TotalTwo jobs$0.22Import step not in this table

How do you place the split cues?

Cue times are measured on the clip you caption, not on the original capture. Trim first with the default exact precision, so the cut starts where you asked and the run begins at 0:00 of the new file. Then take each split from your timer, subtract the in-point you used for the trim, and give each card three seconds on screen.

The example sends three cues in one request. Replace the video URL with the media.sume.com address of your trimmed clip.

import os
import requests

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("Set SUME_API_KEY first")

splits = [("Split 1  0:41.2", 41.2), ("Split 2  1:18.9", 78.9), ("PB  2:03.4", 120.0)]
cues = [{"text": t, "start": s, "end": s + 3.0} for t, s in splits]

resp = requests.post(
    "https://api.sume.com/v1/video-captions",
    headers={"Authorization": f"Bearer {key}", "Idempotency-Key": "pb-splits-001"},
    json={
        "video_url": "https://media.sume.com/artifacts/artf_demo/pb-run.mp4",
        "style": "black-outline",
        "cues": cues,
    },
    timeout=60,
)
resp.raise_for_status()
print(resp.json())

What limits apply to a long capture?

Video trim takes a source of up to 1800 seconds and returns up to 900 seconds, so a half-hour capture is the ceiling for a single cut. If your recording is longer, split it before import. The caption price is stated for videos of at most 60 seconds, so keep a highlight to a minute, which is also a good length for a short-form post.

If you later want a different look, the captions page lets you restyle with a source_caption_id instead of paying to transcribe again. With authored cues there is nothing to transcribe, but the style field and design overrides still let you change colors and placement for a channel look.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume