Which Sume audio endpoint to call: TTS, STT, music, detach, timeline

A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.

5 min readSume
All posts

Match the job to the endpoint: speak text with TTS, transcribe with STT, compose with the music router, pull sound out of a video with audio detach, and join or trim audio with Timeline audio. Everything is an async job with an id, fetched by GET /v1/jobs/:id/result once the status is completed.

The map

Sume audio endpoints, limits and list prices, from the Sume OpenAPI reference and API pricing page, read 2026-10-01.
NeedEndpointKey limitPrice
Speech from text, no engine choicePOST /v1/tts-1.0/generate20,000 characters$0.0475 per 1,000 chars
Speech with a model idPOST /v1/tts-router/generatesonic-3.6 and others$0.0475 per 1,000 chars
Transcript with word timesPOST /v1/stt-1.0/transcribe10 minutes per request$0.01 per minute
MusicPOST /v1/music-router/generateprompt 1 to 5,000 chars$0.125 per generation
Audio out of a videoPOST /v1/audio-detachsource 1,800 s$0.01 per job
Join or split audioPOST /v1/timeline-1.0/audio20 parts, 1,800 s$0.01 per job
Mix voice, bed and videoPOST /v1/timeline-1.0/renderaudio up to 1,800 s$0.10 per output minute

Common chains

  • Dubbing prep: import the video with POST /v1/media-imports, detach its audio, transcribe the result. See extracting audio for transcription.
  • Narrated video: TTS per scene, join the takes with Timeline audio, then render with a music bed. See joining voiceover takes.
  • Captions: transcribe with sentence segmentation and use the segments as cue times.

Rules that trip people up

  • /v1/tts-1.0/generate has no engine picker; the schema refuses model_id. Use the router when you want to choose.
  • Music rejects duration, duration_seconds and a non-empty negative_prompt. Put length in the prompt.
  • Timeline audio requires an Idempotency-Key header, and parts must share a channel layout.
  • Audio detach needs a media.sume.com video, so import first.

Limits

Sume has no live, streaming voice session; every call here is a job. For the cost comparison with live APIs, see realtime voice cost per hour. MCP clients get the same capabilities as tools: tts_create, stt_create, music_create, audio_detach, timeline_audio, jobs_wait and jobs_result. Full reference: docs.sume.com.

More in Developers

All Developers posts

Written by Sume