HeyGen 30-minute avatar video: what Sume does in one request
HeyGen says it can make a 30-minute talking video in one pass. One Sume avatar video request covers 4-60 seconds, so longer pieces are split into jobs.
A single Sume avatar video request accepts a script or scene plan that Sume estimates at 4-60 seconds, so a 30-minute talking video is a series of jobs, not one call. HeyGen's July 2026 update says its avatar can now "generate a single 30-minute AI talking video in one pass"; that is a different limit from Sume's.
What does HeyGen say about 30 minutes?
The July 2026 release notes describe the change as "6x the industry ceiling and 10x our previous maximum". Those are HeyGen's own claims on its own page, and this post does not test them.
What is the limit on a Sume avatar video?
The avatar video docs say scripts and multi-scene plans are accepted when the estimated duration is 4-60 seconds inclusive, and tell you to shorten longer scripts or split them into multiple jobs. The preview docs state the same window for previews.
| Source | Stated limit for one talking video |
|---|---|
| Sume avatar video request | 4-60 seconds, estimated from the script |
| HeyGen (July 2026 notes) | 30 minutes in one pass |
| Sume inline captions | Estimated duration above 60 seconds is rejected |
How do I get 30 minutes out of 60-second jobs?
Cut the script into segments of up to 60 seconds and send each as its own request with the same avatar_handle, so the presenter stays the same across segments. That is 30 requests for 30 minutes at the maximum window. Give every request its own Idempotency-Key.
Join the finished clips in Timeline 1.0 if you need one file; the earlier post on avatar video longer than 60 seconds walks through it.
Will 30 submissions be accepted at once?
Only up to your workspace's accepted job capacity. The generation admission docs list processing concurrency and queue capacity by plan: on Free that is 1 and 5, so 6 accepted jobs; on Pro 4 and 20, so 24. Beyond that, a submit returns 429 queue_full. Submit in batches, or read generation_limits from a submit response.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen Avatar V's 15-second recording vs a Sume photo avatar
HeyGen's Avatar V starts from a 15-second recording. Sume's avatar API starts from a prompt, traits or one public HTTPS image, with no recording step.
- HeyGen Edit Look: retouching an AI avatar, and the Sume route
HeyGen's Edit Look retouches an existing avatar in place. Sume's docs show no such edit, so the route is a new avatar from a retouched photo and a new handle.
- Shortest video an AI dub or face swap accepts: 5 s vs 4 s
Synthesia's dubbing page says a video must be at least 5 seconds. Sume's Beta face swap plans for about 4-15 seconds and avatar videos take 4-60. Side by side.
- Does a face-swapped video need a YouTube AI disclosure?
YouTube asks for disclosure when content makes a real person appear to say or do something they didn't. What that means for a Sume Beta face-swap output.
Written by Sume