Sume TTS takes transcript or transcript_source, never both
A Sume TTS request accepts exactly one text input: literal transcript, or a transcript_source that points at a stored script. How to pick.

Send exactly one text input to Sume TTS: either transcript, the literal text, or transcript_source, a reference to a script that already lives in a thread. Sending both, or neither, is not a valid request. For ordinary API use you will almost always send transcript (up to 20,000 characters). transcript_source exists for flows where the spoken text must come from an accepted script and not from whatever a caller typed.
If you got a validation error on a TTS call, check this first.
What does each input mean?
The Sume OpenAPI spec defines the request with transcript (maximum 20,000 characters) as an alternative to transcript_source. In Sume's repository notes on source-bound TTS, TTS 1.0 and the TTS Router both accept exactly one text input, and the worker receives the resolved transcript either way.
| Input | Use it when | Limit |
|---|---|---|
| transcript | You have the text in hand | 20,000 characters |
| transcript_source | The text must come from a stored, accepted script | Resolved from the source reference |
| Both | Never | Not a valid request |
| Neither | Never | Not a valid request |
When should you reach for transcript_source?
Only when your workflow keeps scripts as records with revisions and you want the audio tied to one of them. Generic chat TTS can still send literal text. Workflows that require a source reference reject literal text even when it is correct, so do not use a source reference as a way around a limit; it is not one.
- Literal transcript: simplest, works for scripts and one-off lines.
- Source reference: stale or mismatched revisions fail rather than speak the wrong text.
- Either way, the 1,200 second audio cap still applies.
What does a valid request look like?
The literal form, with nothing else competing for the text:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
headers=H, timeout=60,
json={
"transcript": "Thanks for calling. How can we help?",
"voice": {"id": os.environ["VOICE_ID"]},
"language": "en",
})
print(r.status_code)
print(r.text[:300])
What about long scripts?
One request, one text input, 20,000 characters. For longer text, split it and join the audio afterwards; the arithmetic is in 100,000 characters over five TTS jobs, and sentence-level output is covered in cutting a voiceover into sentence clips.
Sources
Related posts
More in Developers
- Sume TTS volume 0.5 to 2.0: set the voiceover level before the mix
generation_config volume is a multiplier from 0.5 to 2.0 on Sume TTS. Use it to match narration loudness across jobs before you join or mix them.
- Check for a newer Sume CLI release in CI without auto-upgrading
sume update --check reports whether a newer GitHub Release exists and changes no files. Run it on a schedule, log it, and bump your pinned tag by pull request.
- Sume uploadFile with raw bytes needs a content type, or no request
uploadFile refuses a Uint8Array or ArrayBuffer without contentType and throws SumeUploadError at step create before any call. Pass a typed Blob or a MIME.
- Patching Supabase Postgres 17.11 vs Sume's 10-attempt webhook budget
Supabase's September 25 Postgres 15.19 and 17.11 releases fix 44 CVEs. A restart can outlast Sume's ten 30-second webhook attempts, so plan a redeliver.
Written by Sume