ElevenLabs dubbing 180 min, 3 GB limits vs Sume STT 10-min chunks

ElevenLabs dubbing allows 180 minutes in the app and 3 GB per file via API. Sume STT takes 10 minutes per call, so long videos need a chunk plan.

6 min readSume
All posts

The size limits side by side

The ElevenLabs dubbing page lists automatic dubbing at up to 1 GB and 180 minutes in the app, or 3 GB per file through the API. Dubbing Studio is up to 1 GB and 45 minutes. It also lists three concurrent dubbing jobs on self-serve plans and ten on Enterprise.

Sume has no single dubbing job, so there is no single file limit. Each step has its own: STT takes up to 10 minutes of audio per call (duration_seconds 1-600), TTS renders up to 1,200 seconds from 20,000 characters, audio detach takes a source up to 1,800 seconds and returns at most 900 seconds, and Timeline audio outputs up to 1,800 seconds.

Per-step limits (read 2026-10-03)
StepSume limitWhere it comes from
Extract audioSource 1,800 s, output 900 sAudio detach
TranscribeUp to 10 minutes per callSTT 1.0 duration_seconds 1-600
Speak20,000 characters, 1,200 sTTS 1.0
JoinOutput up to 1,800 s, 20 partsTimeline audio

Planning a 60-minute video

A one-hour video does not fit in one detach call, since output is capped at 900 seconds. Use the range field to take 15-minute windows: four calls cover an hour. Each window is under the STT limit only if you split further, so cut transcription windows of 10 minutes or less.

  • Detach in windows with range, at most 900 s each.
  • Transcribe in pieces of 10 minutes or less.
  • Offset each piece's timestamps by its start time when you merge.
  • Render speech per segment group and join with Timeline audio, whose output tops out at 1,800 s, so assemble long dubs in sections.

Who wins at what size

For one very long file, a packaged product with a 180-minute limit is simpler. For a catalog of many short clips, per-step control is an advantage: each job is retried independently with its own Idempotency-Key, and you can swap a voice without redoing transcription.

Concurrency differs too. The vendor caps concurrent dubbing jobs by plan, while on Sume your limits are the normal rate limits and generation admission rules.

Offsets are the part people get wrong

When you transcribe a long file in pieces, every piece starts its timestamps at zero. Add the piece's start time to every word and segment before you merge, or captions after the first piece will appear at the wrong time.

Check the seam between pieces. Cut at silence if you can, so no word is split across two STT calls.

  • Record each window's start in seconds beside its job id.
  • Merge transcripts in order, offsetting every timestamp.
  • Spot-check one caption in each piece against the video.

Takeaway

Do the math on your longest file first. If it is under about 10 minutes, Sume's steps map one-to-one. If it is an hour, budget windows and offset timestamps, or use the packaged product for that file.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume