Captions drift after cutting pauses: burn them on the final render

Cut pauses first, then caption. Word times from the original clip no longer match once silence is gone. Run Sume video-captions on the trimmed render.

4 min readSume
All posts

Caption after you cut. Every second of pause you remove moves every later word earlier, so word times captured from the original clip are wrong for the trimmed one. Run Sume video-captions on the final trimmed render and let it align to that file (Sume docs: Video captions, read 2026-10-06).

The order is the whole fix. A pipeline that transcribes once, cuts the pauses and then burns the old timings produces captions that drift further behind with each removed gap.

What happens if you reuse the old times

Say a clip has a 1.2 second pause at 10 seconds and a 0.8 second pause at 30 seconds. After both are cut, a word originally at 40 seconds now plays at 38. Captions burned at 40 would appear two seconds late. The error accumulates, so it is invisible early in the video and obvious at the end.

There are two clean routes. Either let the captions job run its own speech-to-text on the trimmed file, or recompute the word times yourself by subtracting the removed pause lengths and send them as words.

Route one: caption the final file

Send the final render's URL to POST /v1/video-captions. The job costs $0.20 for a clip up to 60 seconds (read 2026-10-06 in the Sume docs). Speech-to-text works only when the clip has audible speech; a silent clip fails with caption_no_speech. If you also have the script, send it as script_text and the job keeps the speech-to-text timings as the source of truth and aligns your text to them.

Route two: shift the times yourself

If you already have word times and the list of removed pauses, you can send words with text, start and end. Sume then does not run speech-to-text and burns your text at your times. Send only one of script_text, words, cues and segments. The helper below shifts word times by the cut gaps.

def shift(words, cuts):
    # cuts: list of (cut_start, cut_end) in original time
    out = []
    for w in words:
        removed = sum(min(w["start"], c1) - c0
                      for c0, c1 in cuts if w["start"] > c0)
        removed = max(0.0, removed)
        out.append({"text": w["text"],
                    "start": round(w["start"] - removed, 3),
                    "end": round(w["end"] - removed, 3)})
    return out

words = [{"text": "hello", "start": 1.0, "end": 1.4},
         {"text": "again", "start": 11.5, "end": 11.9}]
print(shift(words, [(10.0, 11.2)]))
Two ways to caption a pause-trimmed video, read 2026-10-06 against Sume docs
RouteYou sendRisk
Caption final fileFinal render URLNeeds audible speech
Send your own wordswords with text, start, endYour shifted times must be right

Check at the end, not the start

Scrub to the last ten seconds of the final file and read the captions against the audio. Drift compounds, so the end shows it first. If it lines up there, it lines up everywhere before it.

Keep the order in your pipeline

Write the order into the pipeline as stages with names: inspect, cut, render, caption, deliver. A stage that reads the output of the one before it cannot be reordered by accident, and a reviewer can see at a glance that captions come last.

If the pipeline is run by an agent, state the order in the instruction too, and ask for the final file's URL to be passed to the captions step. An agent told only to add captions and cut pauses may choose the wrong order, so removing the choice removes the bug.

Know what a restyle does before you reach for it. The Sume docs say a restyle of an existing caption job reuses the word timings that job already has, so speech-to-text does not run a second time, and the price does not change because a restyle is still a render (Video captions, read 2026-10-06). That is convenient for changing a font, and a trap for drift: if the first job was run on the uncut file, a restyle keeps the old timings. After you cut pauses, start a new caption job on the new file instead of restyling the old one.

Silent clips are the other edge. Speech-to-captions needs audible speech and fails with caption_no_speech otherwise, so a trimmed clip that is mostly music will not caption by itself. For that case send cues or segments with text, start and end in seconds, which burns your text at your times without speech-to-text. Whichever route you use, send only one of script_text, words, cues and segments, and remember the clip must be 60 seconds or shorter at $0.20 a job.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume