tts_source_too_large 422: Sume TTS 20,000-character cap on a selection
A transcript_source selection that resolves past 20,000 characters returns 422 tts_source_too_large. Split it by sentence ids and price each job.

tts_source_too_large is a 422 from Sume TTS. It fires when the sentences you selected from an accepted script resolve to a transcript longer than 20,000 characters. The limit is the same as for a literal transcript, which the request schema caps at 20,000 characters. The message reads "Resolved transcript exceeds 20000 characters."
Note which side of the call it fires on. The source API accepts up to 1,000 script_texts blocks of up to 20,000 characters each, so a long script can be accepted whole. The cap bites at generation time, when you ask one job to read too many sentences. Source resolution is thread-scoped: a request with transcript_source but no thread behind it gets 400 tts_text_source_required ("Source resolution requires a thread."). The literal transcript path has no such requirement.
Split on sentence boundaries
Read the manifest first: every sentence has a character_count. Walk the list and start a new job whenever adding the next sentence would pass your budget, keeping each group contiguous. Contiguity matters, because a selection with a gap is refused with tts_sentence_selection_invalid. One detail: when a selection spans two source blocks, Sume joins them with a newline, which is one extra character per block break.
def groups(counts, budget=20000):
out, cur, n = [], [], 0
for i, c in enumerate(counts):
if cur and n + c + 1 > budget:
out.append(cur)
cur, n = [], 0
cur.append(i)
n += c + 1
if cur:
out.append(cur)
return out
counts = [4800, 7200, 6100, 5900, 3000]
print(groups(counts))What the split costs
Sume charges $0.0475 per 1,000 characters, rounded up to the cent per job with a one-cent minimum, so splitting adds at most a cent of rounding per extra job. The table prices a 45,000-character script in the three obvious ways.
| Split | Jobs | Characters per job | Cost |
|---|---|---|---|
| Three equal groups | 3 | 15,000 | $0.72 x 3 = $2.16 |
| 20,000 + 20,000 + 5,000 | 3 | 20,000 / 20,000 / 5,000 | $2.14 |
| Nine chapters of 5,000 | 9 | 5,000 | $2.16 |
Check before you submit
Add the character_count values of your chosen ids (plus one per block break) and compare with 20,000 yourself. That is cheaper than discovering the 422 in a pipeline, although the 422 itself costs nothing, since no job is created.
Sources
Related posts
More in Developers
- tts_text_source_conflict 400: transcript or transcript_source
Sume TTS wants exactly one of transcript and transcript_source. Sending both, neither or a malformed source returns 400. Which code, and how to fix.
- Sume TypeScript SDK createImage: a retired model id fails tsc
Sume's @sume-com/sdk lists accepted image model ids as a string union, so gpt-image-1 fails to compile. Use tsc as the migration checklist.
- TypeScript types for a Sume job status: narrow on sume_status
Type the Sume job envelope as a discriminated union on sume_status, so a switch covers queued to canceled and the compiler flags a missed case. Runs on Node 22.
- Unit test a transcription retry loop with a fake 429 in Python
Test your Sume STT retry code without calling the API: inject the POST and sleep, return a 429 with retry-after, and assert the same key is sent twice.
Written by Sume