AI dubbing API: build the pipeline from STT, TTS and audio joins

Sume has no one-call dubbing route. Chain audio detach, STT 1.0, your translation, TTS 1.0 and Timeline audio into a dub. Routes, limits and the $0.01 steps.

4 min readSume
All posts

Sume has no single dubbing endpoint, so an AI dubbing API on Sume is five steps you chain yourself: detach the audio, transcribe it, translate the sentences in your own code, speak them with TTS 1.0, and join the spoken sentences into one file. The three media steps are priced at $0.01 per detach job, $0.01 per audio minute of speech-to-text and $0.01 per join job; speech is billed by character.

Which call does each stage use?

Each stage is an ordinary job with its own route, so you can retry or swap one without redoing the rest.

Dubbing stages and the Sume route for each, read 2026-09-29.
StageRouteWhat you get back
Extract speechPOST /v1/audio-detachA new audio file; the video is untouched
TranscribePOST /v1/stt-1.0/transcribeText plus words[] timings, optional sentence segments[]
TranslateYour codeSume has no dubbing route
SpeakPOST /v1/tts-1.0/generateAudio, with word timings and sentence segments on request
JoinPOST /v1/timeline-1.0/audioOne gapless file plus segments[] offsets

How do I get clean audio and sentence timings?

Import the video first, because detach only accepts your workspace's media.sume.com video. The docs call sample_rate: 16000 with channels: "mono" the speech-to-text shape. Then send the returned audio_url to STT with sentence segmentation.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dub-detach-001" \
  -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
        "sample_rate": 16000, "channels": "mono" }'

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "audio_url": "<audio_url from the detach result>",
        "language_code": "en",
        "segmentation": { "mode": "sentence" } }'

How do I speak the translated sentences?

Translate outside Sume, keeping each sentence's start and end from STT. Send the translation to POST /v1/tts-1.0/generate with a voice, the target language, timestamps.words: true and segmentation.mode: "sentence". Use a wav container: only wav or raw output returns an audio_url per sentence. Then join the sentence files with operation: "concat"; the result's segments[] give the offsets to line the pictures up against. Set language for every non-English transcript, since it defaults to English.

What should I keep between the steps?

Keep three things in your own store: the job id of every step, the sentence list with its start and end times, and the media.sume.com URL of each finished audio file. Every step is an asynchronous job you poll or receive by webhook, and the API reference says a wait of up to 30 seconds is only a bounded HTTP wait, not a job limit, so submit with async and poll the job status rather than resubmitting a paid job.

Because each stage stores its output as a file, a wrong translation costs you only the speech and join steps again; the extraction and transcript stay as they were. Check the sentence count after translation, since one line per source sentence is what lets the join offsets map back to the original timing.

Where does the pipeline stop?

Limits to plan around: detach takes a source up to 1800 seconds and outputs up to 900, so a longer track needs a range. A concat takes 1 to 20 parts per job, and the joined audio is capped at 1800 seconds. TTS takes up to 20,000 characters per transcript. The result is new speech, not new lip movement; for that see lip sync versus dubbing.

The audio-only path is described in Timeline audio; the join step is not re-synthesis, so the new voice keeps the timing the segments gave it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume