Re-base caption words after a trim: subtract actual_start_seconds
Caption words must use the trimmed file's clock. A word at 12.3-12.7 s becomes 1.9-2.3 s when the trim really started at 10.4, not 0.3-0.7 s.

If you pass your own words to video captions, their start and end are read against the trimmed file, so every STT time from the source clip must be shifted by the trim's real in-point. The trim result reports that in-point as actual_start_seconds; when you asked for 12.0 s with precision: "keyframe" and the cut began at 10.4 s, a word at 12.3-12.7 s belongs at 1.9-2.3 s, and subtracting 12.0 would put it at 0.3-0.7 s, 1.6 s too early.
Why the two numbers differ
Video trim has two precisions. exact, the default, re-encodes and is frame-accurate, so actual_start_seconds equals what you sent. keyframe stream-copies and the cut can start a GOP early, which is why the docs tell you to re-base against actual_start_seconds. The same arithmetic holds either way, so always read the field from the result and never reuse the request value.
| Quantity | Value | How |
|---|---|---|
| Requested start | 12.0 s | what you sent |
| actual_start_seconds | 10.4 s | read from the trim result |
| Word in the source clock | 12.3 to 12.7 s | from STT words[] |
| Correct caption time | 1.9 to 2.3 s | 12.3 - 10.4 and 12.7 - 10.4 |
| Wrong (used 12.0) | 0.3 to 0.7 s | 1.6 s early |
The shift in code
The function drops words that end before the file begins, skips words with no timing (the schema leaves start and end optional), and keeps only words inside the 60-second window that caption timings accept. Send the result as words; you can send only one of script_text, words, cues and segments, and with words Sume does not run speech-to-text.
def rebase(words, requested_start, actual_start, keep_from=0.0):
"""Shift STT word times from the SOURCE clock to the trimmed file.
Use actual_start_seconds from the trim result, not the start you sent."""
out = []
for w in words:
if "start" not in w or "end" not in w:
continue
if w["end"] <= actual_start:
continue # before the file begins
out.append({**w,
"start": round(max(w["start"] - actual_start, 0), 3),
"end": round(w["end"] - actual_start, 3)})
return [w for w in out if w["end"] <= 60] # caption timings cap at 60 sLimits
words,cuesandsegmentstimings are limited to 60 s; the flat $0.20 caption price is stated for videos up to 60 s.- With
exactprecision on anoutputconform, times do not change; only the pixels do. - If you cut the video at several places and join it, shift by each slot's
starttoo, not only the trim in-point. - Caption
video_urlmust be a public HTTPS URL that Sume can fetch; amedia.sume.comartifact from the trim result works.
Sources
Related posts
More in Developers
- A Sume job failed: read GET /v1/jobs/:id/events, then act
The events endpoint is the first stop for a failed or canceled Sume job: the eight event names, what stays hidden, error categories, and a TypeScript reader.
- Read your Sume plan from ratelimit-limit on a GET: 4,800 to 48,000
A GET /v1/balance returns ratelimit-limit. 4800, 12000, 24000 or 48000 reads a minute maps to Free, Pro, Startup or Scale, and a full Wan 720p queue reserve.
- Remote MCP with an API key: x-api-key or Bearer on Sume?
Sume's hosted MCP accepts a key as Authorization: Bearer or as x-api-key. Both see the full tool set; writes and paid calls still need an idempotency_key.
- Replay a sume/auto video submit with one key, assert one job
Submit model sume/auto twice with the same Idempotency-Key and assert the same job id and the same route. Why a replay is safe and why Sume hides the family.
Written by Sume