Dub a 20-minute video: audio detach 900 s cap, STT 600 s, TTS 1,200 s
A 20-minute video needs chunking before a dub: audio detach outputs at most 900 s, STT reservation tops out at 600 s, and TTS fails past 1,200 s.

A 20-minute video is 1,200 seconds, which is over three limits in the Sume dubbing path, so you chunk it. Audio detach accepts a source up to 1,800 seconds but outputs at most 900, so a whole 1,200-second track needs a range. The STT request's duration_seconds hint maxes at 600. And TTS fails with tts_duration_exceeded when the synthesized audio runs past 1,200 seconds. Plan chunks of about 10 minutes.
All limits are from the Sume docs and OpenAPI, read 2026-10-02.
Which limits apply at each step?
Each step of a self-built dub has its own cap. The table lists those that decide the chunk size.
Audio detach and timeline audio cost $0.01 per job each, per their docs; confirm live in GET /v1/catalog.
| Step | Endpoint | Limit |
|---|---|---|
| Pull the audio | POST /v1/audio-detach | Source up to 1,800 s; output up to 900 s; use range past that |
| Transcribe | POST /v1/stt-1.0/transcribe | duration_seconds 1 to 600, used for usage reservation |
| Speak | POST /v1/tts-1.0/generate | Up to 20,000 characters; audio over 1,200 s fails with tts_duration_exceeded |
| Join or split | POST /v1/timeline-1.0/audio | 1 to 20 parts for concat; 1 to 20 ranges for split |
How do I chunk the audio?
Call audio detach twice with range: { "start": 0, "end": 600 } and { "start": 600, "end": 1200 }. Ask for format: wav, channels: mono and sample_rate: 16000, which the docs call the STT shape. The source must be a media.sume.com video in your workspace, so import first with POST /v1/media-imports.
Cut chunks at a pause rather than mid-sentence; transcribe a first pass to find the pauses, or choose ranges at scene breaks you already know.
How do I put the dub back together?
Translate each chunk's text, synthesize it per language with timestamps.words: true and segmentation.mode: "sentence" with a wav output, then join the per-chunk files with operation: "concat" on timeline audio. The docs say the join is sample-domain, with no re-synthesis and no silence at the seams, and the result's segments[] gives the offsets to re-base against.
Translated speech usually runs longer or shorter than the original, so a straight swap drifts against the picture. The sentence-segment guide shows how to keep timing.
What should I do first?
Test the whole chain on one 60-second slice in one language before you chunk the full video, and re-use the detached audio rather than detaching again for a second language.
Sources
Related posts
More in Developers
- Facebook Reels API: start, upload, finish and the video_state values
Publishing a Facebook Reel takes three calls: upload_phase start, a file upload to rupload, then finish with video_state PUBLISHED, SCHEDULED or DRAFT.
- Fix one sentence in an AI avatar video without a full re-render
ElevenLabs can regenerate only edited dubbing regions. A Sume avatar video is one job, so the fix is to keep clips short and join them: the cost math.
- FLUX API 402, 403 and 503 errors, and Sume's equivalents
BFL returns 402 for credits, 403 for key permission, 503 for load. Sume returns 402 insufficient_credits and 503 provider_capacity_exceeded. What to do.
- Gemini API video upload limits vs how Sume takes media inputs
Gemini accepts video inline under 100 MB or through the File API up to 20 GB paid and 2 GB free. How Sume's video routes take URLs, and where they refuse.
Written by Sume