Which API: speech, transcript, music, captions or audio track?
Pick the Sume endpoint by the job: TTS for narration, STT for a transcript, Music for a score, captions for burned text, audio detach for a video's sound.

Match the endpoint to what you hold and what you want. If you hold text and want a voice, use TTS. If you hold audio and want text, use STT. If you hold a video and want its sound, use audio detach. If you hold a finished clip and want words burned on it, use video captions. If you hold nothing and want a score, use Music. All of them share one job lifecycle, so the code to wait for them is the same.
One lifecycle for all five
Every job returns an id on submit. Read GET /v1/jobs/:id/status, and when result_ready is true read /result. Before then /result returns 409 job_not_completed. mode: "sync" waits up to 30 seconds at most, so for longer work use async or a signed webhook_url.
Pick by the thing you hold
Read the first column as your starting point. The input limits are from the Sume docs and the API schema.
| You hold | You want | Endpoint | Input that matters |
|---|---|---|---|
| Text | A voice file | POST /v1/tts-1.0/generate | Transcript up to 20000 characters |
| Audio URL | Text with word timings | POST /v1/stt-1.0/transcribe | Up to 600 s per file |
| Video on media.sume.com | Its audio track | POST /v1/audio-detach | Source up to 1800 s |
| Finished public video URL | Burned captions | POST /v1/video-captions | Speech in the clip, or authored cues |
| A prompt | A music track | POST /v1/music-1.0/generate | Music 1.0 (retiring): prompt up to 5000 characters; Music Router is the recommended route |
Two errors to expect
Two mistakes recur. The first is sending a silent clip to captions: with no audible speech, the job fails with caption_no_speech, and the fix is to send cues with text, start and end. The second is sending a duration to Music 1.0. Music 1.0 has no duration field, so the docs say not to send one; write the length in the prompt.
A short decision rule
Ask what you hold, then what you need. Text in and sound out is TTS. Sound in and text out is STT. A video in and sound out is detach. A finished video in and text on screen is captions. Nothing in and a score out is Music.
If more than one answer fits, do the cheaper and smaller job first. Detach is a flat fee and does no model inference, so it is a good first step. Run the expensive model only on the audio you need, and not on the whole file.
Chain them when you need more. Detach the audio, transcribe it, translate the text yourself, speak it with TTS, join the lines with timeline audio, and render with Timeline 1.0. Each step is a separate job with its own idempotency key. Public rates are listed in GET /v1/catalog, so read the live price there before you build a budget.
Sources
Related posts
More in Developers
- Sume TTS source errors: which are safe to retry (status table)
Every tts_ error code in Sume's script-source API with its HTTP status, whether it charges, and whether a retry can help. A table for client error handling.
- Word timestamps to video frame numbers at 29.97 fps in Python
Sume STT words[] carry start and end in seconds. Convert them to frame indexes with exact 30000/1001 math so cuts do not drift on long timelines.
- Workers OAuth Provider v1 or Sume's hosted MCP: which do I need?
Cloudflare's v1 OAuth library is for building your own MCP server. To call Sume tools from Claude or Cursor, connect Sume's hosted endpoint and skip the build.
- Zed context_servers for the hosted Sume server: no header means OAuth
Add the hosted Sume server to Zed's settings.json context_servers with a url. With no Authorization header Zed runs the MCP OAuth flow, so start with mcp:read.
Written by Sume