Sume STT words[] cap: what words_truncated and words_total mean

Sume STT returns word timings capped at 20,000 entries and sets words_truncated and words_total when it hits the cap. A 600-second job stays far below it.

4 min readSume
All posts

Sume STT returns up to 20,000 word entries per job. If the provider returns more, the result sets words_truncated: true and words_total with the real count, so timings are never dropped silently. Because one job is limited to 600 seconds of audio, you will almost never reach the cap.

All of this is from the Sume OpenAPI contract result schema. The captions behaviour mentioned later is from the Video captions page.

The fields on a finished STT job

A completed STT job exposes text, language_code and language_probability when available, and words[] ordered by start, where each entry has word, start and end in seconds from the audio start, and sometimes a provider type such as word or spacing. With segmentation.mode: "sentence", you also get segments[] with index, text, start, end and duration_seconds, gapless so that each end equals the next start.

STT result fields relevant to the word cap (per Sume OpenAPI, read 2026-10-10)
FieldMeaning
words[]Word timings, ordered by start; present on every STT result, possibly empty
words_truncatedPresent and true only when words was capped
words_totalTotal timed tokens the provider returned; present only with words_truncated
segments[]Sentence segments, only when requested; no sliced audio files

Why the cap rarely matters

Fast speech runs near 3 to 4 words per second, so 600 seconds is about 2,000 words, and with spacing tokens counted as entries perhaps double that. Even four times that is far below 20,000. The cap is a guard against a pathological result, such as a provider returning a token per character, not something you plan around.

I am giving a rough speaking-rate estimate here, not a Sume figure; the documented fact is only the 20,000 cap and the 600-second job limit, and the schema itself says a 600-second transcript stays well under it.

Check for truncation anyway

Handle the flag in code rather than assuming. If words_truncated is true, treat the word list as incomplete, use words_total to see how much is missing, and either re-run on shorter windows or fall back to text for display only. A caption burner that silently stops mid-video is a worse failure than a clear error.

The Video captions page says Sume burns the STT word timings when you do not send your own words, and that sending words skips speech-to-text a second time. That is the other reason to keep the word list: it can feed captions without another STT charge.

A defensive pattern

Read the result, then branch:

  • If words_truncated is absent, use words[] as the full list.
  • If it is true, log words_total and split the source into shorter windows.
  • Always filter entries with type: spacing before building subtitle words.
  • Keep duration_seconds accurate so the reservation matches the audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume