ElevenLabs dubbing 180-minute app limit vs 3 GB API and Sume detach

ElevenLabs dubs up to 180 minutes in the app or 3 GB via API. Sume detach takes 1,800 s of source and 900 s of output per job, so long videos need ranges.

5 min readSume
All posts

ElevenLabs lists its automatic dubbing limits as 1 GB and 180 minutes in the app, or 3 GB per source file through the API. Sume's equivalent step has hard second limits instead: audio detach accepts a source video of up to 1,800 seconds and returns at most 900 seconds of audio per job, so anything longer than 15 minutes needs a range.

The ElevenLabs numbers are from its Dubbing page (read 2026-10-10), which also gives 1 GB and 45 minutes for the editor. The Sume numbers come from Audio detach and Timeline audio.

The limits side by side

Both systems have size or duration ceilings, but they measure different things. ElevenLabs measures the whole source file; Sume measures the audio you extract and the minutes you transcribe. Sume STT jobs also cap at 600 seconds per request, a documented 10-minute maximum, so a long dub is chunked at that step too.

Long-source limits (ElevenLabs read 2026-10-10; Sume per docs.sume.com and OpenAPI)
LimitElevenLabs dubbingSume pipeline
App source1 GB and 180 minutesNot applicable
API source3 GB per fileDetach: source up to 1,800 s
Editor source1 GB and 45 minutesNot applicable
Per-job audio outputNot statedDetach: up to 900 s; STT: duration up to 600 s

Cutting a 40-minute video into ranges

A 40-minute video is 2,400 seconds, above Sume's 1,800-second source cap for detach, so it cannot go through detach in one piece. Video trim has the same 1,800-second source cap, so the practical fix is to export or import the video as shorter files, for example two 20-minute parts, before you start. Each piece then needs detach ranges of at most 900 seconds, and STT in windows of at most 600 seconds.

For a 25-minute video (1,500 seconds), detach with three ranges of 500 seconds each gives three jobs at $0.01 apiece, and each one is also under the 600-second STT cap. The docs say to detach once and split afterwards with timeline audio when you need many ranges; here, three direct range detaches are simpler.

Keeping the pieces in order

Name each piece by its start offset in your own metadata, because Sume returns each file separately and the job envelope does not know your plan. After STT, add each window's start offset to its word times so the transcript is continuous. The STT word timings are seconds from the start of that audio, not of the whole video.

On the output side, timeline audio concat joins up to 20 parts per job and returns segment offsets you can use to place video later. A 40-minute dub can be assembled from three groups of 20 parts and a final join, as long as the produced audio stays within the 1,800-second limit per job; longer than that means two output files.

When the vendor's limit is the better fit

If you have one 150-minute webinar and want a dub in one call, ElevenLabs' 180-minute app limit or 3 GB API limit fits better than chunking by hand. Sume's route is better when you want line-level control and to pay per sentence. Pick based on which of those matters, not on the headline limit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume