ElevenLabs dubbing 180 min, 3 GB limits vs Sume STT 10-min chunks
ElevenLabs dubbing allows 180 minutes in the app and 3 GB per file via API. Sume STT takes 10 minutes per call, so long videos need a chunk plan.

The size limits side by side
The ElevenLabs dubbing page lists automatic dubbing at up to 1 GB and 180 minutes in the app, or 3 GB per file through the API. Dubbing Studio is up to 1 GB and 45 minutes. It also lists three concurrent dubbing jobs on self-serve plans and ten on Enterprise.
Sume has no single dubbing job, so there is no single file limit. Each step has its own: STT takes up to 10 minutes of audio per call (duration_seconds 1-600), TTS renders up to 1,200 seconds from 20,000 characters, audio detach takes a source up to 1,800 seconds and returns at most 900 seconds, and Timeline audio outputs up to 1,800 seconds.
| Step | Sume limit | Where it comes from |
|---|---|---|
| Extract audio | Source 1,800 s, output 900 s | Audio detach |
| Transcribe | Up to 10 minutes per call | STT 1.0 duration_seconds 1-600 |
| Speak | 20,000 characters, 1,200 s | TTS 1.0 |
| Join | Output up to 1,800 s, 20 parts | Timeline audio |
Planning a 60-minute video
A one-hour video does not fit in one detach call, since output is capped at 900 seconds. Use the range field to take 15-minute windows: four calls cover an hour. Each window is under the STT limit only if you split further, so cut transcription windows of 10 minutes or less.
- Detach in windows with
range, at most 900 s each. - Transcribe in pieces of 10 minutes or less.
- Offset each piece's timestamps by its start time when you merge.
- Render speech per segment group and join with Timeline audio, whose output tops out at 1,800 s, so assemble long dubs in sections.
Who wins at what size
For one very long file, a packaged product with a 180-minute limit is simpler. For a catalog of many short clips, per-step control is an advantage: each job is retried independently with its own Idempotency-Key, and you can swap a voice without redoing transcription.
Concurrency differs too. The vendor caps concurrent dubbing jobs by plan, while on Sume your limits are the normal rate limits and generation admission rules.
Offsets are the part people get wrong
When you transcribe a long file in pieces, every piece starts its timestamps at zero. Add the piece's start time to every word and segment before you merge, or captions after the first piece will appear at the wrong time.
Check the seam between pieces. Cut at silence if you can, so no word is split across two STT calls.
- Record each window's start in seconds beside its job id.
- Merge transcripts in order, offsetting every timestamp.
- Spot-check one caption in each piece against the video.
Takeaway
Do the math on your longest file first. If it is under about 10 minutes, Sume's steps map one-to-one. If it is an hour, budget windows and offset timestamps, or use the packaged product for that file.
Sources
Related posts
More in Comparisons
- ElevenLabs dubbing 32 speakers per file vs Sume one voice per TTS job
ElevenLabs dubbing handles up to 32 speakers in a file. Sume composes dubbing from STT and TTS, one voice per job. What that means for multi-speaker video.
- ElevenLabs Dubbing v2 skips the transcript: Sume steps you can read
ElevenLabs says Dubbing v2 conditions on the source performance, not a transcript. Sume dubs run through text you can read and fix before any voice is made.
- ElevenLabs Opus 48 kHz output vs Sume TTS mp3, wav and raw formats
ElevenLabs lists Opus, MP3, PCM and mu-law outputs. Sume TTS 1.0 offers mp3, wav and raw with six sample rates. See which fits your pipeline.
- ElevenLabs previous_text and next_text vs Sume TTS: split scripts
ElevenLabs lets you pass previous_text and next_text for prosody continuity. Sume TTS has no such fields, so here is how to keep long scripts smooth.
Written by Sume