Sume STT words_truncated: the 20,000-word cap and what it means

Sume STT caps words[] at 20,000 entries and sets words_truncated and words_total if that is reached. A 600-second job should stay well under the cap.

3 min readSume
All posts

Sume's STT result caps words[] at 20,000 entries. If the cap is reached, the result also carries words_truncated: true and words_total with the full count. A 600-second job stays well under the cap, so you only see these fields on unusual audio.

What the fields say

The schema states that timings are never dropped silently. words_truncated is present only when the list was capped, and words_total is present only alongside it. The list is ordered by start, and the field is always present, even if empty.

STT result word fields (read 2026-10-03)
FieldTypeWhen present
wordsarray of word, start, end, typeAlways on STT results
words_truncatedbooleanOnly when capped
words_totalintegerOnly with words_truncated
segmentsarrayWhen segmentation.mode is sentence

Why it rarely fires

The job duration limit is 600 seconds. At a fast speaking rate that is a few thousand words, and a token list that includes spacing entries roughly doubles the entries. Even then it sits far below 20,000. A truncation therefore signals something odd, such as a music bed being tokenised.

Handling it in code

Check words_truncated before you trust the list for alignment. If it is true, use segments or split the audio into shorter ranges with audio detach and run each part separately. Never assume a missing flag means a bug; absence means the list is complete.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume