Sume STT words[] cap: what words_truncated and words_total mean
Sume STT returns word timings capped at 20,000 entries and sets words_truncated and words_total when it hits the cap. A 600-second job stays far below it.

Sume STT returns up to 20,000 word entries per job. If the provider returns more, the result sets words_truncated: true and words_total with the real count, so timings are never dropped silently. Because one job is limited to 600 seconds of audio, you will almost never reach the cap.
All of this is from the Sume OpenAPI contract result schema. The captions behaviour mentioned later is from the Video captions page.
The fields on a finished STT job
A completed STT job exposes text, language_code and language_probability when available, and words[] ordered by start, where each entry has word, start and end in seconds from the audio start, and sometimes a provider type such as word or spacing. With segmentation.mode: "sentence", you also get segments[] with index, text, start, end and duration_seconds, gapless so that each end equals the next start.
| Field | Meaning |
|---|---|
| words[] | Word timings, ordered by start; present on every STT result, possibly empty |
| words_truncated | Present and true only when words was capped |
| words_total | Total timed tokens the provider returned; present only with words_truncated |
| segments[] | Sentence segments, only when requested; no sliced audio files |
Why the cap rarely matters
Fast speech runs near 3 to 4 words per second, so 600 seconds is about 2,000 words, and with spacing tokens counted as entries perhaps double that. Even four times that is far below 20,000. The cap is a guard against a pathological result, such as a provider returning a token per character, not something you plan around.
I am giving a rough speaking-rate estimate here, not a Sume figure; the documented fact is only the 20,000 cap and the 600-second job limit, and the schema itself says a 600-second transcript stays well under it.
Check for truncation anyway
Handle the flag in code rather than assuming. If words_truncated is true, treat the word list as incomplete, use words_total to see how much is missing, and either re-run on shorter windows or fall back to text for display only. A caption burner that silently stops mid-video is a worse failure than a clear error.
The Video captions page says Sume burns the STT word timings when you do not send your own words, and that sending words skips speech-to-text a second time. That is the other reason to keep the word list: it can feed captions without another STT charge.
A defensive pattern
Read the result, then branch:
- If
words_truncatedis absent, usewords[]as the full list. - If it is true, log
words_totaland split the source into shorter windows. - Always filter entries with
type: spacingbefore building subtitle words. - Keep
duration_secondsaccurate so the reservation matches the audio.
Sources
Related posts
More in Developers
- POST /v1/trending-research: Reels and TikTok niche search over REST
One call returns up to 24 ranked public short videos for a niche from Instagram and TikTok at $0.10 per accepted search. Fields, limits and a Python example.
- Sume TTS 400 tts_language_script_mismatch: Korean encoding check
Sume TTS rejects language ko when the transcript has no Hangul syllable: 400 tts_language_script_mismatch, no charge. Usually the text was decoded wrongly.
- Sume waitForJob in TypeScript: ms timeout, failed jobs resolve
waitForJob in @sume-com/sdk takes milliseconds, resolves for failed and canceled jobs, and throws only on timeout or a failed read. 25-line sample.
- Sume Send test vs Redeliver: which to use for avatar video jobs
Send test posts a dummy webhook.test payload to a URL you type; Redeliver re-sends a real job's terminal event with a fresh signature. When to use each.
Written by Sume