Translate a long transcript in 60-second chunks for captions
Index-Translate's FP8 card sets a 4096-token limit and Sume caption jobs cover 60 seconds. Split cues into windows, translate each, and shift times to zero.

Split a long transcript into windows of cues that each end inside 60 seconds, translate one window per request, and shift each window's times so it starts at zero. Two limits push you there. The FP8 checkpoint of Index-Translate-35B-A3B lists a maximum model length of 4096 tokens, and Sume's caption endpoint only accepts cue times within 0 to 60 seconds.
The two limits are unrelated, but they happen to line up well: a minute of speech is a few hundred words, far below 4096 tokens, so a 60-second window is a safe unit for both.
Which limits are you fitting inside?
The README gives 32768 tokens as the default serving context for the 2B and 9B models, so the 4096 figure is specific to the FP8 build of the 35B-A3B preview. Check which checkpoint you actually call before you decide how much to send.
| Limit | Value | Source |
|---|---|---|
| 35B-A3B FP8 max model length | 4096 tokens | Hugging Face model card |
| 2B and 9B default serving context | 32768 tokens | GitHub README |
| Caption cue times | 0 to 60 seconds | Sume video captions docs |
| Cues per caption request | 200 maximum | Sume video captions docs |
| Characters per cue | 400 maximum | Sume video captions docs |
How do you cut the windows?
Walk the cues in order and start a new window whenever the next cue would end more than 60 seconds after the window's first cue began. Never split a cue, and prefer to break at a long pause so a sentence is not cut across two windows. The function below returns each window with times already shifted to zero.
def windows(cues, span=60.0):
out, cur, base = [], [], None
for c in cues:
if base is None:
base = c["start"]
if cur and c["end"] - base > span:
out.append(cur)
cur, base = [], c["start"]
cur.append({"text": c["text"],
"start": round(c["start"] - base, 3),
"end": round(c["end"] - base, 3)})
if cur:
out.append(cur)
return out
cues = [{"text": "a", "start": 0, "end": 30},
{"text": "b", "start": 31, "end": 59},
{"text": "c", "start": 60, "end": 90}]
for w in windows(cues):
print(w)What happens on the video side?
A caption job burns cues onto one video, and cue times must fall inside the first 60 seconds of it. For a longer video, cut it into 60-second pieces first, one per window, and caption each piece. Video inspect can give you sentence boundaries to cut at, which beats a blind 60-second slice that lands mid-word.
A standalone caption job is priced at $0.20 for videos up to 60 seconds under the current estimate, so a ten-minute video is roughly ten jobs. Re-check the amount against the catalog before a large batch, and stitch the captioned pieces back together afterwards.
What should you keep consistent across windows?
Each window is validated as in the JSON round-trip check, then posted with a public HTTPS video_url for its piece.
- Send the same glossary and the same style instruction with every window.
- Include the last cue of the previous window as read-only context if a sentence crosses the boundary.
- Use the same caption style on every piece so the cut points are not visible.
- Log the window index with every translated cue so a failed window can be retried alone.
Sources
Related posts
More in Developers
- OS credential store vs sume login: where the key lives
Inngest v1.45.0 stores CLI OAuth credentials in the OS credential store. The Sume CLI stores its login key in ~/.sume-com/config.json; use env keys in CI.
- Instagram 4:5 on GPT Image 2.5: ask for 1088x1360, not 1080x1350
GPT Image 2.5 needs both edges as multiples of 16. 1080x1350 fails that rule; 1088x1360 keeps 4:5 and passes. How to ask on Sume and trim to 1080x1350.
- Instagram Reels API: a 100-posts-per-24-hours publish budget
The Instagram content publishing API limits an account to 100 API-published posts per moving 24 hours. A tested Python queue that spreads batch output under it.
- IPv6-only webhook endpoint: test Sume delivery before launch
OpenAI's API now accepts IPv6 connections. If your webhook host is IPv6-only, prove Sume can reach it with POST /v1/webhooks/test-deliveries first.
Written by Sume