TTS 400: Provide exactly one of transcript_source or transcript
This 400 means the TTS body had both transcript and transcript_source, or neither. Send exactly one, plus a voice, and re-run.

The error text is Provide exactly one of transcript_source or transcript. It means your TTS request body carried both fields, or neither. Remove one so exactly one remains, and the request validates.
Why it happens
- A template sends transcript with an empty string and also a transcript_source.
- A script sets transcript_source for one path and forgets to delete the default transcript.
- Neither is set because a variable was undefined.
A body that passes
Use transcript for inline text. The voice goes in voice.id. The price is $0.0475 per 1,000 characters and the cap is 20,000 characters per request.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-001" \
-d '{
"transcript": "Thanks for calling. Your order has shipped.",
"voice": {"id": "YOUR_VOICE_ID"}
}'Cost of the example
The transcript above is 44 characters, so it is 44 x $0.0475 / 1,000 = $0.00209.
| Item | Value |
|---|---|
| Price | $0.0475 per 1,000 characters |
| Max per request | 20,000 characters |
| Default output | mp3, 44,100 Hz, 128,000 bps |
| Text format | Plain transcript, no SSML |
Check before sending
- Exactly one of transcript or transcript_source.
- A voice selector: voice.id, or avatar_id or avatar_handle at the top level.
- Under 20,000 characters.
- A fresh Idempotency-Key per distinct job.
Sources
Related posts
More in Developers
- TTS 400 asking for voice.id or avatar_id / avatar_handle: the fix
The TTS call needs a voice: set voice.id, or a top-level avatar_id or avatar_handle. Without one the request is rejected before any audio is made.
- List Sume TTS Router models before you hardcode a Sonic id
GET /v1/tts-router/models returns the catalog of pass-through TTS models. Read it at startup instead of pinning an id that may change.
- Voice replication API audit checklist before you switch
Gemini 3.8 Flash TTS is GA with voice replication and 150+ voices. Before switching providers, audit these items against Sume's live catalog.
- Reuse TTS word timestamps as caption words, skip a second STT
You already know what the voice said and when. Feed the TTS word timings to the caption job as `words` so brand names are never misheard by speech-to-text.
Written by Sume