Captions out of sync with the audio: check STT word times and offsets

Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.

5 min readSume
All posts

When captions drift away from the speech, the first suspect is a missing time offset. Sume STT returns word times as seconds from the start of the audio you submitted, not from the start of your original file. If you transcribe a slice that began at 600 seconds and then burn its words onto the full video without adding 600, every caption lands ten minutes early. The fix is one addition per word before you send the times anywhere.

Where word times start from

The STT result's words[] entries have start and end, documented as offsets in seconds from the audio start, and optional segments[] use the same basis. For a video inspect transcript the audio is the clip itself, so the times line up with the clip. For audio-detach with a range, the detached file starts at the range start, so its first word is near zero whatever the range was.

That is the whole model: every result is relative to the file you handed to speech-to-text.

Which offset to add, by how the audio was made (Sume docs and schema read 2026-10-07)
How the audio was producedTime zero isAdd this to word times
Whole clip, video inspect with transcribeStart of the clipNothing
Audio detach with no rangeStart of the videoNothing
Audio detach with range start 600600 s into the video600
Timeline audio split starting at 300300 s into the source audio300
Clip later trimmed to start at 4 sOriginal startSubtract 4

Merge slices with the offset applied

Keep the offset next to each result. Then filter the spacing tokens out, because words[] can include entries with type spacing, add the offset, round to milliseconds, and sort by start. The result is a single list you can post to the caption endpoint as words, remembering the limits: at most 1,200 words, each time at most 60 seconds, and the standalone job only handles a 60 second clip.

def merge(results):
    """results: [(offset_seconds, stt_result), ...] -> words for video-captions."""
    words = []
    for offset, result in results:
        for w in result["words"]:
            if w.get("type", "word") != "word":
                continue
            words.append({
                "text": w["word"],
                "start": round(w["start"] + offset, 3),
                "end": round(w["end"] + offset, 3),
            })
    return sorted(words, key=lambda w: w["start"])


part1 = {"words": [{"word": "Hello", "start": 0.1, "end": 0.4, "type": "word"},
                   {"word": " ", "start": 0.4, "end": 0.5, "type": "spacing"}]}
part2 = {"words": [{"word": "again", "start": 0.2, "end": 0.6, "type": "word"}]}
print(merge([(0, part1), (600, part2)]))

Other causes to rule out

If the offsets are right and captions are still early, look at the video. A caption job burns text at the times from speech-to-text, or at the times you send, so if you sent words or cues yourself, the problem is in your times. If you used the default flow and the speech starts late, a silent intro is not a bug: the first word begins when the voice does.

Trimming after transcription is the other common one. If you detached audio, transcribed, then cut the first 4 seconds off the video, subtract 4 from every time or caption the trimmed file instead. And a 2x playback check matters if you plan to publish to platforms that offer faster playback; see the 2x playback timing check.

A quick way to see the drift

Pick a word from the middle of the transcript, note its start, and pull the video frame at that time with POST /v1/video-frames and at: [start]. If the frame shows the speaker with an open mouth on the word, the time is right. If the mouth is closed and the speaker is mid-phrase elsewhere, you have an offset. A frame extract is billed by its Modal compute, so check two or three words and not fifty.

Check the start, the middle and the end of the clip. A constant error means a missing offset; an error that grows across the clip points to a different source of the times, for example words from one render combined with a video from another.

A short checklist

Record the offset of each slice when you create it, not afterward. Store the job id with the offset so you can rebuild the merge. Round times to milliseconds, since the caption schema takes numbers and floating point noise helps nobody. Keep every result unchanged and apply offsets in a merge step, so you can rerun it if a time turns out wrong. And after any trim, recompute: a trimmed video is a new timeline, and the old word times belong to the old one.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume