Which Sume speech endpoint: TTS 1.0, TTS Router, STT or Music?
A decision table for Sume's four audio endpoints: TTS 1.0, TTS Router, STT 1.0 and Music Router. Inputs, limits, ids and what each does not do.

Use POST /v1/tts-1.0/generate for a managed voice, POST /v1/tts-router/generate when you must name a Sonic model, POST /v1/stt-1.0/transcribe to turn a recording into text, and POST /v1/music-router/generate for a music bed. All four are async jobs with the same status and result URLs. None streams.
The four endpoints
| Endpoint | Input | Key limits | Output |
|---|---|---|---|
| TTS 1.0 (sume/tts-1.0) | transcript or transcript_source, plus an avatar or voice id | 1 to 20,000 characters; audio over 1,200 s fails | Hosted audio, optional word timings and sentence segments |
| TTS Router | Same, plus a required model | Sonic ids only: sonic-3.6, 3.5, 3, latest, preview | Same as TTS 1.0; job.model is the id you sent |
| STT 1.0 (sume/stt-1.0) | A public HTTPS audio_url, optional language_code | duration_seconds 1 to 600; word timings always returned | text, words[], optional sentence segments |
| Music Router | prompt, optional image_url, optional model | 1 to 5,000 characters; no duration field | Hosted audio artifact, fixed Music price |
Pick by the question you are asking
- "I want a voice reading my script." TTS 1.0. It always runs
sume/tts-1.0and has no engine picker; sendingmodelreturns 400. - "I need a specific Sonic version, pinned." TTS Router, with a
modelfromGET /v1/tts-router/models. - "I have a recording and need words and times." STT 1.0. Omit
language_codefor auto-detect. - "I need a track under a scene." Music Router, with
sume/music-auto,lyria-3.5orlyria-3-pro.
What none of them do
The TTS Router docs list its non-goals: Eleven and OpenAI engines, streaming TTS, and a routing preset. STT is a file job, not a live stream. Music takes no duration or negative prompt, so length and exclusions go in the brief.
STT can appear on the dev API before the production OpenAPI snapshot lists it, so confirm the path on the current reference before you build on it.
The same shape for all four
Every one returns a job envelope. Poll status_url until result_ready is true, then read result_url. mode: sync waits at most 30 seconds; a timed-out wait means keep polling, not resubmit. Send an Idempotency-Key on every submit.
Over MCP the same families are tts_create, stt_create and music_create, with tts_source_get and tts_source_verify_spine as free reads.
curl https://api.sume.com/v1/jobs/job_123/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/job_123/result \
-H "Authorization: Bearer $SUME_API_KEY"Sources
Related posts
More in Developers
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
- Spend caps for unattended AI agents: how Sume bounds each run
An unattended agent has no one to approve spend, so Sume caps generation per run: required on Agent Completions, and up to $500 on Format runs.
Written by Sume