Dub a 20-minute video: audio detach 900 s cap, STT 600 s, TTS 1,200 s

A 20-minute video needs chunking before a dub: audio detach outputs at most 900 s, STT reservation tops out at 600 s, and TTS fails past 1,200 s.

5 min readSume
All posts

A 20-minute video is 1,200 seconds, which is over three limits in the Sume dubbing path, so you chunk it. Audio detach accepts a source up to 1,800 seconds but outputs at most 900, so a whole 1,200-second track needs a range. The STT request's duration_seconds hint maxes at 600. And TTS fails with tts_duration_exceeded when the synthesized audio runs past 1,200 seconds. Plan chunks of about 10 minutes.

All limits are from the Sume docs and OpenAPI, read 2026-10-02.

Which limits apply at each step?

Each step of a self-built dub has its own cap. The table lists those that decide the chunk size.

Audio detach and timeline audio cost $0.01 per job each, per their docs; confirm live in GET /v1/catalog.

Limits from the Sume audio detach and timeline audio docs and the OpenAPI schemas, read 2026-10-02.
StepEndpointLimit
Pull the audioPOST /v1/audio-detachSource up to 1,800 s; output up to 900 s; use range past that
TranscribePOST /v1/stt-1.0/transcribeduration_seconds 1 to 600, used for usage reservation
SpeakPOST /v1/tts-1.0/generateUp to 20,000 characters; audio over 1,200 s fails with tts_duration_exceeded
Join or splitPOST /v1/timeline-1.0/audio1 to 20 parts for concat; 1 to 20 ranges for split

How do I chunk the audio?

Call audio detach twice with range: { "start": 0, "end": 600 } and { "start": 600, "end": 1200 }. Ask for format: wav, channels: mono and sample_rate: 16000, which the docs call the STT shape. The source must be a media.sume.com video in your workspace, so import first with POST /v1/media-imports.

Cut chunks at a pause rather than mid-sentence; transcribe a first pass to find the pauses, or choose ranges at scene breaks you already know.

How do I put the dub back together?

Translate each chunk's text, synthesize it per language with timestamps.words: true and segmentation.mode: "sentence" with a wav output, then join the per-chunk files with operation: "concat" on timeline audio. The docs say the join is sample-domain, with no re-synthesis and no silence at the seams, and the result's segments[] gives the offsets to re-base against.

Translated speech usually runs longer or shorter than the original, so a straight swap drifts against the picture. The sentence-segment guide shows how to keep timing.

What should I do first?

Test the whole chain on one 60-second slice in one language before you chunk the full video, and re-use the detached audio rather than detaching again for a second language.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume