Synthesia dubbing file size limit: 5 GB or 2.5 hours
Synthesia dubbing accepts uploads up to 5 GB or 2.5 hours, 4K, in .mp4, .webm or .mov. Sume's limits sit at the audio stage: a 10 MB audio_url and 300 s.

Synthesia's dubbing page says an uploaded file can be up to 5 GB or 2.5 hours, at up to 4K (3840x2160), in .mp4, .webm or .mov. A YouTube link has no duration limit. Sume does not publish a whole-video dubbing upload limit; the limits that matter in a Sume pipeline are per stage, such as a 10 MB audio_url and a 300 second duration_seconds on the Fabric endpoint.
Synthesia numbers are from its dubbing page, read 2026-10-01. Sume numbers are from the OpenAPI document and Avatar videos.
What are the Synthesia dubbing limits?
The upload row also lists frame rates of 23.98 to 60 fps and audio sample rates of 8 to 96 kHz. The Dubbing page can dub up to 10 videos in bulk. Under advanced options, a duration setting of Adaptive (the default) changes playback speed to fit the translation, while Original keeps the video speed and adjusts only the voiceover.
| Where | Limit |
|---|---|
| Synthesia upload | Up to 5 GB or 2.5 hours, up to 4K |
| Synthesia YouTube link | No duration limit |
| Synthesia bulk dubbing | Up to 10 videos |
Sume Fabric audio_url | Sume-hosted public HTTPS URL, max 10 MB |
Sume Fabric duration_seconds | 1 to 300 |
| Sume Avatar Video | Estimated duration 4-60 seconds |
Why are the numbers so different?
They bound different things. Synthesia's numbers cap the source video you hand over. Sume's Fabric endpoint takes a still image plus an audio file, and the OpenAPI schema says non-Sume hosts are rejected for audio_url, so audio must already live on the Sume media host. Compare against a TTS segment, not a feature film.
Can Sume put new speech on an existing video?
The models docs say video models do not lip-sync to generated TTS or to a later voice-over, and that every on-camera speaking shot is Fabric with an accepted still plus TTS. So a Sume talking face is generated from a still and audio, not re-voiced from a source video. For a speech-only pipeline, see build an AI dubbing pipeline with STT, translate and TTS.
What should I do with a long source?
Cut the audio into segments that each stay under 10 MB and 300 seconds, upload each to the Sume media host, and make one request per segment. If you only need a Synthesia-sized file dubbed as is, that is a question for Synthesia's page, not Sume's.
Sources
Related posts
More in Developers
- Synthesia Interactive Avatar and LiveKit: Sume's job-based pieces
Synthesia's Interactive Avatar API runs on a LiveKit plugin with your own LLM and STT. Sume offers job pieces: TTS, speech to text and still-plus-audio clips.
- Synthesia rejected status: moderation, error, and your handler
Synthesia marks a moderated video rejected, separate from error. Sume reports content_policy_rejected as a public_reason; an unfunded run is a 402 instead.
- Synthesia video.completed webhook: the download URL is time-limited
Synthesia sends video.completed and video.failed, and its download URL is time-limited. Sume sends job.completed, job.failed and job.canceled with durable URLs.
- Take It Down Act 48-hour removal for AI video apps and URLs
The FTC enforces a 48-hour removal duty on covered platforms. If your AI video app serves generated MP4s from durable public URLs, here is what to plan for.
Written by Sume