Text to speech API voice id: list avatars where voice.status is ready
Sume TTS rejects voice names from other services before any credit is held. Use avatar_id or avatar_handle, or a UUID or voi_ id copied verbatim.
On Sume, find a TTS voice by listing avatars with GET /v1/avatar-1.0/avatars and picking one whose voice.status is ready, then pass its avatar_id or avatar_handle on the TTS request. Sume resolves that avatar's voice at submit time. If you already hold a voice id, send it as voice.id; it must be a UUID (8-4-4-4-12 hex) or a Voices library id (voi_ plus 32 hex).
Anything else, such as a voice name from another TTS service, is rejected synchronously with 400 invalid_voice_id before a job is queued or credits are reserved. This is from the TTS request contract in the repository (read 2026-10-09).
The three ways to select
The contract lists the avatar reference as the discoverable selector: a caller that holds only the API contract never needs a voice id it cannot obtain. avatar_id and avatar_handle are top-level request fields, not properties of voice. If you send both, they must resolve to the same avatar.
| Selector | Where | Rule |
|---|---|---|
| avatar_id | top-level field | From GET /v1/avatar-1.0/avatars; voice.status must be ready |
| avatar_handle | top-level field | Same avatar grammar as the Avatar APIs |
| voice.id | voice object, mode id | UUID or voi_ plus 32 hex; copy verbatim |
| avatar plus voice.id | both | voice.id must equal the avatar's resolved voice id, else 400 |
What a first request looks like
List the avatars, choose one, and send transcript plus the handle. The transcript is required unless you use a transcript_source, and exactly one of the two applies. The maximum is 20,000 characters, at $0.0475 per 1,000 characters.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-first-line-001" \
-d '{"transcript": "Hello from Sume.", "avatar_handle": "studio_presenter"}'Common mistakes
Do not send model or model_id to TTS 1.0; it has no engine picker and returns 400. Use the TTS Router at POST /v1/tts-router/generate with a catalog model id if you want to choose the engine. Do not send provider credentials; authenticate with your Sume API key only.
The default output is mp3 at 44,100 Hz and 128 kbps. Ask for wav with pcm_s16le at 44,100 if the audio goes to a timeline or an avatar mux step.
Reading the result
A finished TTS job records the voice it used as { "mode": "id", "id": "..." }, along with model_id, language, output_format, generation_config and speed. Copy the voice object from a job you liked into the next request to keep the same voice.
Check voice.status before you build a batch. An avatar whose voice is not ready cannot supply a voice, and the error comes back before any credit is held, which is cheaper than finding out halfway through a run.
If your integration is an agent, the hosted MCP tool is tts_create (paid, with an idempotency_key), and tts_source_get and tts_source_verify_spine are free reads for scripts that use transcript_source.
Keep the voice selector in one place in your code. When a voice changes, you update one constant instead of hunting through request builders, and the job records tell you which voice each file used.
Sources
Related posts
More in Developers
- Three blind retries of a 30 s Seedance 720p submit can reserve $52.00
Without an Idempotency-Key, a timeout and two retries on one 30 s Seedance 2.5 clip can book $52.002. The same loop with one key, in curl, plus Wan 3.0 totals.
- Total usage.cost for Sume video job ids with curl, jq and awk
One shell pipeline: read job ids from a file, GET /v1/videos/{id} for each, keep completed jobs, and sum usage.cost with awk. No Python or Node needed.
- Transcribe a 15-second voice note in one request: STT sync mode
Sume STT can answer in the same HTTP call: mode sync with wait_timeout_seconds up to 30 returns 200, otherwise 202 and a poll. A $0.01 minimum, with curl.
- Transcribe the first 60 seconds of 500 support recordings
500 screen-recorded support sessions: detach seconds 0 to 60 ($5.00), then STT 60 s each ($5.00). Total $10.00 at $0.01 per job and per minute.
Written by Sume