LLM-written cue sheet: validate the times before you burn captions

Gemini 3.8 Flash or any LLM can draft caption cues as JSON. Check order, overlap and length in Python, then burn them with /v1/video-captions for $0.20.

5 min readSume
All posts

A language model can draft the caption cues for a silent or lightly narrated episode, but you should treat its output as untrusted input: parse it, check that the times are ordered and inside the clip, and cap each line's length before you send it to Sume. Sume's POST /v1/video-captions then burns exactly the cues you give it, with no speech-to-text, for $0.20 on a video up to 60 seconds (read 2026-10-03).

Google's model page lists gemini-3.8-flash as stable, described as its most intelligent Flash model, with 3.7 and 3.6 Flash labelled previous-generation (read 2026-10-03). The page does not discuss caption timing, and nothing below depends on that model. Any LLM that returns JSON can draft the sheet. What matters is the gate between its answer and the paid call.

What the caption route accepts

cues (also spelled segments) are phrase-level overlay cards, each with text, start and end in seconds. Either form skips speech-to-text and burns exactly that copy at those times. They are exclusive with script_text and words: send one of the four, never two. This is also the route for clips with no speech, because speech-based captions on a silent clip fail with caption_no_speech.

The route does not fix your timing. If a cue starts after the clip ends or two cards overlap, the job may still be billed, so the validation belongs on your side.

Cue sheet rules to enforce before the call (read 2026-10-03)
RuleWhy it matters
start < end for every cueA zero or negative card never shows
Cues sorted by start, no overlapOverlapping cards fight for the same screen space
Last end at or before the clip durationA cue past the end is wasted copy
Text length capped per cardLong cards wrap past a vertical frame
Only one of cues, words, script_textThe fields are mutually exclusive
Clip up to 60 seconds for the $0.20 estimateLonger clips: read the live catalog price

A validator and the call

The script below asks nothing of the model. It takes whatever JSON the model returned, drops anything unsafe, and posts the rest. Set your own character limit per card; 42 is a starting point for a vertical frame, not a Sume rule.

A failed check prints the reason and exits before any billable call, which is the point of the gate.

import json, os, sys, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MAX_CHARS, CLIP_SECONDS = 42, 45.0

def validate(cues):
    prev_end = 0.0
    for i, c in enumerate(cues):
        if not (0 <= c["start"] < c["end"] <= CLIP_SECONDS):
            return f"cue {i}: bad times"
        if c["start"] < prev_end:
            return f"cue {i}: overlaps the previous cue"
        if not c["text"].strip() or len(c["text"]) > MAX_CHARS:
            return f"cue {i}: empty or over {MAX_CHARS} characters"
        prev_end = c["end"]
    return None

cues = json.load(open("llm_cues.json"))
err = validate(cues)
if err:
    sys.exit(err)
r = requests.post(f"{API}/v1/video-captions",
    headers={**H, "Idempotency-Key": "ep6-cues-v1"},
    json={"video_url": os.environ["EPISODE_URL"], "style": "slam", "cues": cues})
print(r.status_code, r.json().get("request_id"))

What to do with the model's draft

Ask for JSON only, and if the model wraps it in prose, strip everything outside the first and last bracket before parsing. Give the model the clip length and the character limit in the prompt, so the common failures are rare. Then validate anyway. Models sometimes return cues in the wrong order, drift past the end, or add a trailing comma that breaks the parse, and each of those costs a render if it reaches Sume.

Keep the sheet in a file beside the episode. A restyle or a corrected word is then a new caption job on the same clean clip with a new idempotency key, for another $0.20, rather than a new generation. Name keys by version, like ep6-cues-v2, because reusing a key with a different body is a conflict.

Read the cues out loud against the picture once. A validator proves the times are legal, not that the words land on the right shot. For a clip with few cards, a thirty-second check by a person is cheaper than a second render.

A worked example: a 45-second silent episode with six cards. Six cues of about 30 characters each is 180 characters of copy, and the validator passes when the first card starts at 1.0 seconds, each next card starts at or after the previous end, and the last ends at 43.5. The caption job is $0.20 whatever the number of cards, so there is no saving in cutting cards, only in cutting passes. Draft, validate, burn once, and keep the sheet for the next episode as a template whose text you replace.

If the clip has speech, you may not need cues at all. Pass script_text instead and Sume keeps its speech-to-text word timings as the timing source and aligns your wording to them. That path can fail with script_alignment_mismatch, so keep cues for clips without speech or where you want to dictate the timing.

What Sume does and does not do

Sume burns the cues you send, in the style you name, and returns a captioned video. It does not write the copy, does not repair a bad cue sheet, and does not guarantee a card fits the frame if you ignore the length limit. Drafting and validation stay in your code.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume