STT words into caption words: rename word to text, stay under 60 s

Sume STT returns words as {word, start, end}; the caption job wants {text, start, end}, end above start, 60 s or less, 1200 words at most. Map and filter.

4 min readSume
All posts

Sume STT 1.0 returns words[] as {word, start, end}, but the caption job's words input is {text, start, end}, so you rename the key and drop any row that would fail validation. The caption schema wants end greater than start, both times at most 60 seconds, text of 1 to 200 characters, and 1 to 1200 words in total. Sending words skips speech-to-text, so you pay for the transcript once.

The two shapes side by side

The STT result documents words[] as word-level timings in seconds from the start of the audio, always present on a completed job and possibly empty. The caption request describes words as timed wording that skips STT and is mutually exclusive with script_text, cues and segments.

The difference is small and easy to miss: one key name, plus stricter limits on the caption side.

STT result words vs caption request words (Sume docs and request schemas, repo read 2026-10-05)
PropertySTT 1.0 resultCaption request
Text keywordtext
Time keysstart, end (seconds)start, end (seconds, at most 60)
Zero-length tokenPossibleRejected: end must exceed start
Maximum length20000 entries, then words_truncated1200 entries
Empty listPossibleRejected: at least 1

A mapping that passes validation

This function renames the key, trims the text, drops rows that would fail and stops at 1200. It reads a result dictionary, so you can test it on a saved response.

def to_caption_words(stt_result, limit=60.0):
    out = []
    for w in stt_result.get("words", []):
        text = str(w.get("word", "")).strip()[:200]
        start, end = w.get("start"), w.get("end")
        if not text or start is None or end is None:
            continue
        if end <= start or end > limit:
            continue
        out.append({"text": text, "start": start, "end": end})
    return out[:1200]

print(to_caption_words({"words": [
    {"word": "Hello", "start": 0.0, "end": 0.4},
    {"word": "  ", "start": 0.4, "end": 0.5},
    {"word": "world.", "start": 0.5, "end": 0.5},
]}))

Limits

A mapping like this is a filter, and it has a cost: dropping a word silently removes it from the burned captions, so log what you drop. If a clip is longer than 60 seconds, the standalone caption job is not the right tool; see the video captions docs for the ceiling and the price, which is $0.20 per accepted job under the current estimate. When you only need to fix one word, send source_caption_id and the corrected words instead of redoing the transcript.

Checks worth running before you submit

The caption schema validates every word, so a bad row rejects the whole request. A pre-flight pass is cheaper than a failed submit.

  • Every text must be 1 to 200 characters after trimming, so a whitespace-only token has to go.
  • end must be strictly greater than start; an STT token that is zero length fails the rule.
  • Both times must be 60 seconds or less, the same ceiling as the caption job.
  • The list holds 1 to 1200 words, so a long clip needs chunking by the minute.
  • Never send script_text or cues alongside it; the request treats them as conflicting sources.
  • Keep the original STT job id next to the caption job, so a missing word can be traced back to the transcript rather than to your filter.
  • If STT returned a token such as a lone punctuation mark, expect it to be dropped or merged; the filter treats it like any other row.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume