Captions out of sync with the audio: check STT word times and offsets
Captions running early or late usually trace to an unapplied offset. How Sume STT word times work, which offset to add, and a Python merge that applies it.

When captions drift away from the speech, the first suspect is a missing time offset. Sume STT returns word times as seconds from the start of the audio you submitted, not from the start of your original file. If you transcribe a slice that began at 600 seconds and then burn its words onto the full video without adding 600, every caption lands ten minutes early. The fix is one addition per word before you send the times anywhere.
Where word times start from
The STT result's words[] entries have start and end, documented as offsets in seconds from the audio start, and optional segments[] use the same basis. For a video inspect transcript the audio is the clip itself, so the times line up with the clip. For audio-detach with a range, the detached file starts at the range start, so its first word is near zero whatever the range was.
That is the whole model: every result is relative to the file you handed to speech-to-text.
| How the audio was produced | Time zero is | Add this to word times |
|---|---|---|
| Whole clip, video inspect with transcribe | Start of the clip | Nothing |
| Audio detach with no range | Start of the video | Nothing |
| Audio detach with range start 600 | 600 s into the video | 600 |
| Timeline audio split starting at 300 | 300 s into the source audio | 300 |
| Clip later trimmed to start at 4 s | Original start | Subtract 4 |
Merge slices with the offset applied
Keep the offset next to each result. Then filter the spacing tokens out, because words[] can include entries with type spacing, add the offset, round to milliseconds, and sort by start. The result is a single list you can post to the caption endpoint as words, remembering the limits: at most 1,200 words, each time at most 60 seconds, and the standalone job only handles a 60 second clip.
def merge(results):
"""results: [(offset_seconds, stt_result), ...] -> words for video-captions."""
words = []
for offset, result in results:
for w in result["words"]:
if w.get("type", "word") != "word":
continue
words.append({
"text": w["word"],
"start": round(w["start"] + offset, 3),
"end": round(w["end"] + offset, 3),
})
return sorted(words, key=lambda w: w["start"])
part1 = {"words": [{"word": "Hello", "start": 0.1, "end": 0.4, "type": "word"},
{"word": " ", "start": 0.4, "end": 0.5, "type": "spacing"}]}
part2 = {"words": [{"word": "again", "start": 0.2, "end": 0.6, "type": "word"}]}
print(merge([(0, part1), (600, part2)]))Other causes to rule out
If the offsets are right and captions are still early, look at the video. A caption job burns text at the times from speech-to-text, or at the times you send, so if you sent words or cues yourself, the problem is in your times. If you used the default flow and the speech starts late, a silent intro is not a bug: the first word begins when the voice does.
Trimming after transcription is the other common one. If you detached audio, transcribed, then cut the first 4 seconds off the video, subtract 4 from every time or caption the trimmed file instead. And a 2x playback check matters if you plan to publish to platforms that offer faster playback; see the 2x playback timing check.
A quick way to see the drift
Pick a word from the middle of the transcript, note its start, and pull the video frame at that time with POST /v1/video-frames and at: [start]. If the frame shows the speaker with an open mouth on the word, the time is right. If the mouth is closed and the speaker is mid-phrase elsewhere, you have an offset. A frame extract is billed by its Modal compute, so check two or three words and not fifty.
Check the start, the middle and the end of the clip. A constant error means a missing offset; an error that grows across the clip points to a different source of the times, for example words from one render combined with a video from another.
A short checklist
Record the offset of each slice when you create it, not afterward. Store the job id with the offset so you can rebuild the merge. Round times to milliseconds, since the caption schema takes numbers and floating point noise helps nobody. Keep every result unchanged and apply offsets in a merge step, so you can rerun it if a time turns out wrong. And after any trim, recompute: a trimmed video is a new timeline, and the old word times belong to the old one.
Sources
Related posts
More in Developers
- Connect a new MCP client to Sume: five calls that prove it works
After you add https://mcp.sume.com/mcp to a new client, run mcp_health, tools_list, tools_schema, account_me and catalog_list. What each result should show.
- Convert an SRT file to Sume caption cues in Python
Sume captions take no SRT upload, but cues carry the same start, end and text. A 26-line Python script turns an SRT into cues and posts them for $0.20.
- Cost per ad variant: build a ledger from usage.cost on Sume
Every completed /v1/videos poll carries usage.cost. Sum it by hook and ending to get the cost per ad variant before media spend. Node script and the caveats.
- Create an AI avatar and its first talking video in one bash script
Two Sume jobs in order: create the avatar, wait, then render a talking video with its handle. A bash script with curl and jq, plus the cost of both steps.
Written by Sume