Which Sume audio endpoint do I need? TTS, STT, music, detach, split

Sume has separate endpoints for text to speech, speech to text, music, detaching video audio, and splitting or joining audio. Here is which to call for what.

5 min readSume
All posts

Sume splits audio work across six endpoints: /v1/tts-1.0/generate makes speech from text, /v1/stt-1.0/transcribe turns speech into text with word times, /v1/music-router/generate makes music, /v1/audio-detach pulls the audio track out of a video, and /v1/timeline-1.0/audio joins or splits audio files. Pick by what you have and what you need.

Paths come from the OpenAPI document behind the API reference and the Sume API reference (read 2026-10-06).

Which one for which job?

All of them are jobs: submit, poll GET /v1/jobs/{id}/status, read /result.

Sume audio endpoints (read 2026-10-06)
You haveYou needCall
A scriptA voiceover filePOST /v1/tts-1.0/generate
A script and a model choiceA voiceover on a pinned Sonic modelPOST /v1/tts-router/generate
A recorded audio URLText and word timesPOST /v1/stt-1.0/transcribe
A promptA music trackPOST /v1/music-router/generate
A video on media.sume.comIts audio as WAV or MP3POST /v1/audio-detach
Several audio filesOne gapless filePOST /v1/timeline-1.0/audio with concat
One audio fileSeveral slicesPOST /v1/timeline-1.0/audio with split

What are the limits to remember?

TTS takes up to 20,000 characters per request and fails if the synthesized audio runs over 1,200 seconds. STT takes a duration_seconds of 1 to 600. Detach caps the source at 1,800 seconds and the output at 900. Timeline audio takes 1 to 20 parts or ranges per job.

What is not in the list?

There is no streaming TTS and no streaming speech to text on Sume. The TTS Router carries Cartesia Sonic models only, so other TTS families are not offered.

Sources

More in Developers

All Developers posts

Written by Sume