TTS from an accepted script: transcript_source instead of pasted text
Sume TTS can read an accepted script by script_revision_id and sentence_ids, not pasted text. MCP tools tts_source_get and tts_source_verify_spine support it.

Instead of pasting text into transcript, a Sume TTS request can reference a server-held accepted script with transcript_source: a script_revision_id plus a list of sentence_ids. The API resolves the text itself, so the audio always matches the approved wording. The MCP tools tts_source_get and tts_source_verify_spine are the read side of the same flow.
This matters when someone approves a script and a different person, or an agent, generates the audio: nobody can slip in altered text.
What the contract says
In the OpenAPI contract, transcript_source references accepted server-owned script text. You obtain the ids from GET /v1/tts-1.0/source in the same authenticated session. sentence_ids must be unique, contiguous and in source order, from 1 to 1000 ids. The API resolves the text before preview, usage reservation or dispatch. It accepts no client text, paths, hashes or receipts. That is the guarantee: the request can only point at text, not smuggle it in.
The inline alternative is the transcript field, up to 20,000 characters. You use one or the other.
The MCP side
The MCP tools doc lists tts_create among the paid tools, with an idempotency_key, and describes two free read tools. tts_source_get returns the accepted-script manifest to use with tts_create and transcript_source. tts_source_verify_spine checks selected TTS jobs against the accepted script, so you can confirm that the audio you are about to use really came from the approved text.
| Tool | Kind | Purpose |
|---|---|---|
| tts_create | Paid, idempotency_key | Create a TTS job, with transcript or transcript_source |
| tts_source_get | Read, free | Manifest of the accepted script |
| tts_source_verify_spine | Read, free | Check selected TTS jobs against the accepted script |
| stt_create | Paid, idempotency_key | Create a transcription job |
| music_create | Paid, idempotency_key | Create a music job |
A review workflow
Once a script has been accepted, read the manifest with tts_source_get, then call tts_create with the revision and the sentences you want. When the jobs complete, run tts_source_verify_spine over them. A pass means the audio belongs to the approved script; a failure means somebody generated from something else.
- Use the same idempotency_key on retries so a timeout never double-bills.
- Select contiguous sentences only; gaps are rejected.
- Keep revisions immutable in your own records.
- Treat verification as a gate before captions or publishing.
When to use plain transcript
For a one-off voiceover you wrote yourself, the transcript field is simpler. Reach for transcript_source when approval and generation are separate steps, or when several people or agents touch the pipeline.
Related posts
More in Developers
- TTS volume 0.5 to 2: set narration gain before mixing with music
Sume TTS 1.0 generation_config.volume runs from 0.5 to 2 alongside speed 0.6 to 1.5. How to set narration level before you mix with a music bed.
- TTS mp3 bit_rate vs wav: fit a voiceover under the 10 MB Fabric limit
Sume TTS mp3 bit rates run 32k to 192k. At 128k a 300-second voiceover is about 4.8 MB, under the 10 MB Fabric audio limit. Mono 16 kHz wav is 9.6 MB.
- TTS word timings to burned-in captions: send them as words on Sume
Sume's TTS can return word start and end times; the caption job accepts words with text, start and end and skips transcription. How to wire them together.
- Two workers, one order: Idempotency-Key from order id and version
Two queue workers pick up the same order and both submit to Sume. Build the key from order id plus version so duplicates collapse and edits still create a run.
Written by Sume