Which Sume audio endpoint do I need? TTS, STT, music, detach, split
Sume has separate endpoints for text to speech, speech to text, music, detaching video audio, and splitting or joining audio. Here is which to call for what.

Sume splits audio work across six endpoints: /v1/tts-1.0/generate makes speech from text, /v1/stt-1.0/transcribe turns speech into text with word times, /v1/music-router/generate makes music, /v1/audio-detach pulls the audio track out of a video, and /v1/timeline-1.0/audio joins or splits audio files. Pick by what you have and what you need.
Paths come from the OpenAPI document behind the API reference and the Sume API reference (read 2026-10-06).
Which one for which job?
All of them are jobs: submit, poll GET /v1/jobs/{id}/status, read /result.
| You have | You need | Call |
|---|---|---|
| A script | A voiceover file | POST /v1/tts-1.0/generate |
| A script and a model choice | A voiceover on a pinned Sonic model | POST /v1/tts-router/generate |
| A recorded audio URL | Text and word times | POST /v1/stt-1.0/transcribe |
| A prompt | A music track | POST /v1/music-router/generate |
| A video on media.sume.com | Its audio as WAV or MP3 | POST /v1/audio-detach |
| Several audio files | One gapless file | POST /v1/timeline-1.0/audio with concat |
| One audio file | Several slices | POST /v1/timeline-1.0/audio with split |
What are the limits to remember?
TTS takes up to 20,000 characters per request and fails if the synthesized audio runs over 1,200 seconds. STT takes a duration_seconds of 1 to 600. Detach caps the source at 1,800 seconds and the output at 900. Timeline audio takes 1 to 20 parts or ranges per job.
What is not in the list?
There is no streaming TTS and no streaming speech to text on Sume. The TTS Router carries Cartesia Sonic models only, so other TTS families are not offered.
Sources
More in Developers
- Which API: speech, transcript, music, captions or audio track?
Pick the Sume endpoint by the job: TTS for narration, STT for a transcript, Music for a score, captions for burned text, audio detach for a video's sound.
- Sume TTS source errors: which are safe to retry (status table)
Every tts_ error code in Sume's script-source API with its HTTP status, whether it charges, and whether a retry can help. A table for client error handling.
- Which Sume video models take 9:16? Query the catalog, do not guess
GET /v1/videos/models lists supported_aspect_ratios and supported_durations per model. A jq filter picks the vertical models that reach your platform's length.
- Which wait to use for a 30 second clip: sync, webhook or a poll
A table of the waits Sume documents for a 30 s clip: a bounded sync of up to 30 s, webhooks with retries, and a roughly 20 minute client deadline. Which fits.
Written by Sume