Which API: speech, transcript, music, captions or audio track?

Pick the Sume endpoint by the job: TTS for narration, STT for a transcript, Music for a score, captions for burned text, audio detach for a video's sound.

4 min readSume
All posts

Match the endpoint to what you hold and what you want. If you hold text and want a voice, use TTS. If you hold audio and want text, use STT. If you hold a video and want its sound, use audio detach. If you hold a finished clip and want words burned on it, use video captions. If you hold nothing and want a score, use Music. All of them share one job lifecycle, so the code to wait for them is the same.

One lifecycle for all five

Every job returns an id on submit. Read GET /v1/jobs/:id/status, and when result_ready is true read /result. Before then /result returns 409 job_not_completed. mode: "sync" waits up to 30 seconds at most, so for longer work use async or a signed webhook_url.

Pick by the thing you hold

Read the first column as your starting point. The input limits are from the Sume docs and the API schema.

Speech and music endpoints, read 2026-10-06
You holdYou wantEndpointInput that matters
TextA voice filePOST /v1/tts-1.0/generateTranscript up to 20000 characters
Audio URLText with word timingsPOST /v1/stt-1.0/transcribeUp to 600 s per file
Video on media.sume.comIts audio trackPOST /v1/audio-detachSource up to 1800 s
Finished public video URLBurned captionsPOST /v1/video-captionsSpeech in the clip, or authored cues
A promptA music trackPOST /v1/music-1.0/generateMusic 1.0 (retiring): prompt up to 5000 characters; Music Router is the recommended route

Two errors to expect

Two mistakes recur. The first is sending a silent clip to captions: with no audible speech, the job fails with caption_no_speech, and the fix is to send cues with text, start and end. The second is sending a duration to Music 1.0. Music 1.0 has no duration field, so the docs say not to send one; write the length in the prompt.

A short decision rule

Ask what you hold, then what you need. Text in and sound out is TTS. Sound in and text out is STT. A video in and sound out is detach. A finished video in and text on screen is captions. Nothing in and a score out is Music.

If more than one answer fits, do the cheaper and smaller job first. Detach is a flat fee and does no model inference, so it is a good first step. Run the expensive model only on the audio you need, and not on the whole file.

Chain them when you need more. Detach the audio, transcribe it, translate the text yourself, speak it with TTS, join the lines with timeline audio, and render with Timeline 1.0. Each step is a separate job with its own idempotency key. Public rates are listed in GET /v1/catalog, so read the live price there before you build a budget.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume