Sume TTS takes transcript or transcript_source, never both

A Sume TTS request accepts exactly one text input: literal transcript, or a transcript_source that points at a stored script. How to pick.

5 min readSume
All posts

Send exactly one text input to Sume TTS: either transcript, the literal text, or transcript_source, a reference to a script that already lives in a thread. Sending both, or neither, is not a valid request. For ordinary API use you will almost always send transcript (up to 20,000 characters). transcript_source exists for flows where the spoken text must come from an accepted script and not from whatever a caller typed.

If you got a validation error on a TTS call, check this first.

What does each input mean?

The Sume OpenAPI spec defines the request with transcript (maximum 20,000 characters) as an alternative to transcript_source. In Sume's repository notes on source-bound TTS, TTS 1.0 and the TTS Router both accept exactly one text input, and the worker receives the resolved transcript either way.

Text input choices, Sume spec and repo docs read 2026-10-04
InputUse it whenLimit
transcriptYou have the text in hand20,000 characters
transcript_sourceThe text must come from a stored, accepted scriptResolved from the source reference
BothNeverNot a valid request
NeitherNeverNot a valid request

When should you reach for transcript_source?

Only when your workflow keeps scripts as records with revisions and you want the audio tied to one of them. Generic chat TTS can still send literal text. Workflows that require a source reference reject literal text even when it is correct, so do not use a source reference as a way around a limit; it is not one.

  • Literal transcript: simplest, works for scripts and one-off lines.
  • Source reference: stale or mismatched revisions fail rather than speak the wrong text.
  • Either way, the 1,200 second audio cap still applies.

What does a valid request look like?

The literal form, with nothing else competing for the text:

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
    headers=H, timeout=60,
    json={
        "transcript": "Thanks for calling. How can we help?",
        "voice": {"id": os.environ["VOICE_ID"]},
        "language": "en",
    })
print(r.status_code)
print(r.text[:300])

What about long scripts?

One request, one text input, 20,000 characters. For longer text, split it and join the audio afterwards; the arithmetic is in 100,000 characters over five TTS jobs, and sentence-level output is covered in cutting a voiceover into sentence clips.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume