Which Sume audio endpoint to call: TTS, STT, music, detach, timeline
A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.

Match the job to the endpoint: speak text with TTS, transcribe with STT, compose with the music router, pull sound out of a video with audio detach, and join or trim audio with Timeline audio. Everything is an async job with an id, fetched by GET /v1/jobs/:id/result once the status is completed.
The map
| Need | Endpoint | Key limit | Price |
|---|---|---|---|
| Speech from text, no engine choice | POST /v1/tts-1.0/generate | 20,000 characters | $0.0475 per 1,000 chars |
| Speech with a model id | POST /v1/tts-router/generate | sonic-3.6 and others | $0.0475 per 1,000 chars |
| Transcript with word times | POST /v1/stt-1.0/transcribe | 10 minutes per request | $0.01 per minute |
| Music | POST /v1/music-router/generate | prompt 1 to 5,000 chars | $0.125 per generation |
| Audio out of a video | POST /v1/audio-detach | source 1,800 s | $0.01 per job |
| Join or split audio | POST /v1/timeline-1.0/audio | 20 parts, 1,800 s | $0.01 per job |
| Mix voice, bed and video | POST /v1/timeline-1.0/render | audio up to 1,800 s | $0.10 per output minute |
Common chains
- Dubbing prep: import the video with
POST /v1/media-imports, detach its audio, transcribe the result. See extracting audio for transcription. - Narrated video: TTS per scene, join the takes with Timeline audio, then render with a music bed. See joining voiceover takes.
- Captions: transcribe with sentence segmentation and use the segments as cue times.
Rules that trip people up
/v1/tts-1.0/generatehas no engine picker; the schema refusesmodel_id. Use the router when you want to choose.- Music rejects
duration,duration_secondsand a non-emptynegative_prompt. Put length in the prompt. - Timeline audio requires an Idempotency-Key header, and parts must share a channel layout.
- Audio detach needs a media.sume.com video, so import first.
Limits
Sume has no live, streaming voice session; every call here is a job. For the cost comparison with live APIs, see realtime voice cost per hour. MCP clients get the same capabilities as tools: tts_create, stt_create, music_create, audio_detach, timeline_audio, jobs_wait and jobs_result. Full reference: docs.sume.com.
More in Developers
- Sume video tools: public URL or media import first? Per tool
Video captions takes a public HTTPS URL; trim, filter, inspect, frames, compose and detach need a workspace media.sume.com clip. A tool-by-tool input guide.
- Which voice does my avatar speak with? Check voice.status is ready
Sume TTS speaks in an avatar's voice when voice.status is ready. List avatars, check voice.status, then send avatar_id or avatar_handle on the TTS request.
- YouTube captions.insert: 100 MB, 400 quota units, and an SRT build
YouTube captions.insert costs 400 quota units and takes a 100 MB file. Sume returns words and segments, not SRT, so here is the 20-line conversion to upload.
- YouTube videos.insert: 256 GB limit and containsSyntheticMedia
The YouTube Data API videos.insert method takes files up to 256 GB and a status.containsSyntheticMedia flag. A request body for an AI-made upload, with checks.
Written by Sume