How long can HeyGen videos be? Limits by plan and API
HeyGen videos can run 1 minute on Free, 30 on Creator and Pro, 60 on Business, with no max on Enterprise; an API avatar request caps at 30 minutes.

A HeyGen video can be up to 1 minute long on the Free plan, 30 minutes on Creator and Pro, and 60 minutes on Business, and Enterprise has no video duration maximum, according to HeyGen's pricing page. Through HeyGen's API, a single avatar request is one scene capped at 30 minutes.
HeyGen and Synthesia figures come from their own pricing page, API usage limits and video creation docs, read on 2026-09-28. Sume figures come from Generate avatar video and Timeline 1.0.
What is the HeyGen video length limit on each plan?
HeyGen's plan comparison lists two limits: a maximum duration per video, and a separate maximum for videos made with its Avatar IV engine.
| HeyGen plan | Max duration per video | Avatar IV max per video |
|---|---|---|
| Free | 1 minute | 1 minute |
| Creator | 30 minutes | 30 minutes |
| Pro | 30 minutes | 30 minutes |
| Business | 60 minutes | 30 minutes |
| Enterprise | No video duration max | 30 minutes |
How long can a HeyGen API video be?
The API counts length per scene. HeyGen's usage limits give a maximum of 30 minutes per scene and 50 scenes per video:
- An
avatarorimagerequest toPOST /v3/videosrenders a single scene, so 30 minutes is also its whole-video maximum. - Multi-scene studio videos and templates can run as long as the sum of their scenes, as long as each scene stays within 30 minutes.
- A scene over the cap fails during rendering.
- API video is priced per minute and charged by the actual seconds generated. HeyGen API pricing has the rates.
How long can a Synthesia video be?
Synthesia sets its limit by scenes: a video can have up to 150 scenes, each up to 5 minutes long, with a total duration of up to 4 hours. The script box allows 5 minutes of script per scene, and an uploaded voiceover can be at most 5 minutes per scene.
What you can make each month is a separate limit: Synthesia's self-serve plans count usage in credits, and each second of video uses 2 credits.
How long can a Sume avatar video be?
Sume's avatar API is short-form per job. POST /v1/avatar-1.0/talking-video accepts a script or multi-scene plan only when Sume estimates the video at 4-60 seconds inclusive; the docs say to shorten longer scripts or split them into multiple jobs.
- Longer videos are several jobs joined in one Timeline 1.0 render, whose output can run 1-1,800 seconds. Making an AI avatar video longer than 60 seconds walks through the split and the join.
- Lip sync to a separate speech track, such as a TTS file on the Sume media host, is another route: VEED Fabric 1.0 takes
duration_secondsof up to 300 and an audio file of up to 10 MB per request. - Each talking video is billed per second of video, by quality tier.
Which limit matters for my video?
- For one continuous presenter take, the per-scene cap is the number to check: 30 minutes on HeyGen's API, 5 minutes per scene on Synthesia, 60 seconds per job on Sume.
- For a long finished video, check the whole-video cap on your plan, then the monthly allowance or per-minute price that pays for it.
- For short clips (ads, answers, lesson steps), every product above fits, and the price per second or per minute is the number to compare.
Sources
- HeyGen pricing (read 2026-09-28)
- HeyGen API usage limits (read 2026-09-28)
- HeyGen API pricing explained (read 2026-09-28)
- Synthesia video creation (read 2026-09-28)
- Synthesia script box (read 2026-09-28)
- Synthesia: what are credits on self-serve plans (read 2026-09-28)
- Generate avatar video
- Timeline 1.0
- Sume API reference
- API reference
Related posts
More in Sume Avatar 1.0
- AI avatar prompts: how to describe a talking presenter
An AI avatar prompt for video describes one person: age, look, hair, clothing, and expression. What to write, what Sume adds, and what it can't set.
- Kling API avatar: one image plus audio, 2 to 300 seconds
Kling's API has an Avatar endpoint: one reference image plus a 2–300 second audio track becomes a talking video, billed per second. Inputs and Sume options.
- Difference between lip sync and dubbing: which do you need?
Dubbing replaces the speech in a video; lip sync matches a mouth to the audio. Lip-sync dubbing does both. What each means, and which one you need.
- How long should a microlearning video be? Length and AI
No fixed standard: one objective per video, and research on course videos favors 6 minutes or less. How to size lessons and make them with an AI avatar.
Written by Sume