Captions drift after cutting pauses: burn them on the final render
Cut pauses first, then caption. Word times from the original clip no longer match once silence is gone. Run Sume video-captions on the trimmed render.

Caption after you cut. Every second of pause you remove moves every later word earlier, so word times captured from the original clip are wrong for the trimmed one. Run Sume video-captions on the final trimmed render and let it align to that file (Sume docs: Video captions, read 2026-10-06).
The order is the whole fix. A pipeline that transcribes once, cuts the pauses and then burns the old timings produces captions that drift further behind with each removed gap.
What happens if you reuse the old times
Say a clip has a 1.2 second pause at 10 seconds and a 0.8 second pause at 30 seconds. After both are cut, a word originally at 40 seconds now plays at 38. Captions burned at 40 would appear two seconds late. The error accumulates, so it is invisible early in the video and obvious at the end.
There are two clean routes. Either let the captions job run its own speech-to-text on the trimmed file, or recompute the word times yourself by subtracting the removed pause lengths and send them as words.
Route one: caption the final file
Send the final render's URL to POST /v1/video-captions. The job costs $0.20 for a clip up to 60 seconds (read 2026-10-06 in the Sume docs). Speech-to-text works only when the clip has audible speech; a silent clip fails with caption_no_speech. If you also have the script, send it as script_text and the job keeps the speech-to-text timings as the source of truth and aligns your text to them.
Route two: shift the times yourself
If you already have word times and the list of removed pauses, you can send words with text, start and end. Sume then does not run speech-to-text and burns your text at your times. Send only one of script_text, words, cues and segments. The helper below shifts word times by the cut gaps.
def shift(words, cuts):
# cuts: list of (cut_start, cut_end) in original time
out = []
for w in words:
removed = sum(min(w["start"], c1) - c0
for c0, c1 in cuts if w["start"] > c0)
removed = max(0.0, removed)
out.append({"text": w["text"],
"start": round(w["start"] - removed, 3),
"end": round(w["end"] - removed, 3)})
return out
words = [{"text": "hello", "start": 1.0, "end": 1.4},
{"text": "again", "start": 11.5, "end": 11.9}]
print(shift(words, [(10.0, 11.2)]))| Route | You send | Risk |
|---|---|---|
| Caption final file | Final render URL | Needs audible speech |
| Send your own words | words with text, start, end | Your shifted times must be right |
Check at the end, not the start
Scrub to the last ten seconds of the final file and read the captions against the audio. Drift compounds, so the end shows it first. If it lines up there, it lines up everywhere before it.
Keep the order in your pipeline
Write the order into the pipeline as stages with names: inspect, cut, render, caption, deliver. A stage that reads the output of the one before it cannot be reordered by accident, and a reviewer can see at a glance that captions come last.
If the pipeline is run by an agent, state the order in the instruction too, and ask for the final file's URL to be passed to the captions step. An agent told only to add captions and cut pauses may choose the wrong order, so removing the choice removes the bug.
Know what a restyle does before you reach for it. The Sume docs say a restyle of an existing caption job reuses the word timings that job already has, so speech-to-text does not run a second time, and the price does not change because a restyle is still a render (Video captions, read 2026-10-06). That is convenient for changing a font, and a trap for drift: if the first job was run on the uncut file, a restyle keeps the old timings. After you cut pauses, start a new caption job on the new file instead of restyling the old one.
Silent clips are the other edge. Speech-to-captions needs audible speech and fails with caption_no_speech otherwise, so a trimmed clip that is mostly music will not caption by itself. For that case send cues or segments with text, start and end in seconds, which burns your text at your times without speech-to-text. Whichever route you use, send only one of script_text, words, cues and segments, and remember the clip must be 60 seconds or shorter at $0.20 a job.
Sources
Related posts
More in Agents
- Create a Sume schedule by API? Dashboard first, then trigger by API
The Sume API cannot create or edit schedules. Create it in the dashboard with a cron or an api trigger, then call it from code. What each trigger type means.
- Cut a video to 15 seconds for Reddit Engaged Video Views
Reddit began a 15-second Engaged Video Views beta in September. Cut any hosted clip to 15 seconds with Sume video-trim for a flat $0.02, exact or keyframe.
- Keyframe or exact video trim for pause cuts: read actual_start_seconds
Exact trim cuts where you say; keyframe trim is faster but snaps to a keyframe. For pause cuts that decides whether words get clipped. How to choose and check.
- Lower the spend cap for one scheduled Sume run: overrides only go down
A per-run generation_spend_cap_usd on a scheduled run is clamped to the schedule's own cap. You can lower it for one run but never raise it. How it works.
Written by Sume