Caption cue limits: 400 characters, 200 cues, 60 seconds

A caption cue takes 1-400 characters, start of 0 or more, end above start and 60 s or less; a request holds 1-200 cues. Validate locally before the $0.20 job.

4 min readSume
All posts

The cues input to Sume's video captions endpoint takes between 1 and 200 cues, and each cue has text of 1 to 400 characters, a start of 0 or more and an end greater than start, with both times at most 60 seconds. A cue can hold two lines joined by a newline. segments is an alias for the same shape, and either one skips speech-to-text, which also makes it the way to caption a silent clip.

The limits as the schema states them

The request schema describes a cue as a phrase-level overlay: the same timings as words, but the text can be a full card line, or two lines joined by \n, rather than a single spoken token. The worker burns each cue as one on-screen span without running speech-to-text.

Because no STT runs, a clip without any speech does not trigger caption_no_speech when you send cues. The video captions docs list that error with the next action use_overlay_captions.

Cue limits (Sume request schema and docs, repo read 2026-10-05)
FieldLimitNote
cues or segments1 to 200 itemsAlias for each other; use one
text1 to 400 charactersTwo lines with a newline
start0 to 60 secondsSeconds from the video start
endAbove start, at most 60Equal times are rejected
Price$0.20 per accepted jobVideos up to 60 seconds, current estimate

A local validator

Run this before you submit. It checks the same rules and tells you which cue failed, so you fix it before a paid job.

def check_cues(cues):
    errors = []
    if not 1 <= len(cues) <= 200:
        errors.append(f"need 1-200 cues, got {len(cues)}")
    for i, c in enumerate(cues):
        text = str(c.get("text", "")).strip()
        s, e = c.get("start"), c.get("end")
        if not 1 <= len(text) <= 400:
            errors.append(f"cue {i}: text length {len(text)}")
        if s is None or e is None or s < 0 or e > 60 or e <= s:
            errors.append(f"cue {i}: bad times {s}-{e}")
    return errors

print(check_cues([{"text": "Sale ends Friday", "start": 0, "end": 2.5},
                  {"text": "", "start": 3, "end": 3}]))

Limits

Cues are your timings, so Sume does not check that the text matches any speech. The 60 second ceiling belongs to the standalone caption job; longer videos need another route, and a Hangul cue still needs a Hangul style. For a fix to one cue after a burn, restyle or correct through source_caption_id instead of resubmitting everything.

Common ways to trip it

Most rejected cue lists fail for the same few reasons, and all are cheap to check in advance.

  • A cue whose end equals its start, which happens when you round times to a coarse step.
  • Times that run past 60 seconds because the cue list was made from a longer cut of the video.
  • A cue with only whitespace, which is trimmed to nothing and fails the minimum length.
  • More than 200 cues because every word became its own cue; use words for that instead.
  • Sending cues together with segments or script_text, which is a separate conflict.

Where cues come from

Most cue lists start as something else: a translated script, an SRT you cannot upload, or sentence segments from an STT job. The caption endpoint does not accept SRT upload, so convert the file into the cue shape yourself, one object per subtitle with text, start and end. Split any subtitle over 400 characters, and trim any end time that runs past 60 seconds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume