HeyGen Video API 5-15 s at 768p vs Sume avatar video 4-60 s
HeyGen's new Video API makes 5-15 second clips at 480p or 768p from a prompt. Sume's avatar video renders scripts of 4-60 seconds at 720p. Which fits your job?
HeyGen's new generative Video API makes 5-15 second clips at 480p or 768p from a text prompt, an image or references. Sume's Avatar 1.0 talking video is a different product: a stored avatar speaks a script, the estimated length must land between 4 and 60 seconds, and output is 720p.
So the choice is not about which number is bigger. One tool generates scenes from a prompt; the other makes a specific presenter say specific words.
What does HeyGen's Video API accept?
The changelog entry describes POST /v3/models/videos and GET /v3/models/videos/{video_id} for the heygen-video-1 model, with three modes: text_to_video, image_to_video and reference_to_video. Output is 5 to 15 seconds at 480p or 768p, with prompts up to 5,000 characters. Image-to-video follows the proportions of the source image. Reference-to-video accepts up to nine images, three videos and three audio recordings, twelve references in total. Creation returns HTTP 202 with a video id you poll.
What does Sume's avatar video accept?
POST /v1/avatar-1.0/talking-video takes a ready avatar_handle and exactly one of script or video_inputs. Scripts and multi-scene plans are accepted when Sume estimates the video at 4-60 seconds inclusive. aspect_ratio supports 1:1, 3:4, 9:16, 4:3 and 16:9 (default 9:16), and resolution is currently 720p. Quality is standard, plus (the default) or max.
Optional product_image and scene inputs, either a prompt or a public photo URL, add context. The avatar speaks; there is no free-form prompt describing the whole film.
| HeyGen Video API | Sume avatar video | |
|---|---|---|
| Input | Prompt, image, or references | Avatar handle plus script or video_inputs |
| Length | 5-15 seconds | 4-60 seconds, estimated |
| Resolution | 480p or 768p | 720p |
| Who is on screen | Whatever the prompt generates | One resolved avatar per video |
| Returns | 202 and a video id to poll | 202 with job id plus status, result and events URLs |
When do you pick which?
Pick a prompt-driven generator when you want scenery, motion or a stylised shot with no fixed presenter. Pick the avatar route when the output is a person delivering a script: product pitches, explainers, onboarding, localised greetings.
If you need non-avatar clips on Sume, the Video page covers the video models, which are separate from Avatar 1.0 and have their own limits. Do not expect the avatar route to take a long prompt and invent a scene.
How do you get past 60 seconds on Sume?
You cannot stretch one avatar video beyond the window: longer scripts must be shortened or split into multiple jobs. A common pattern is one job per chapter, then sequencing the finished clips. Each job is its own paid generation with its own reservation, so budget the sum.
Inline captions are also capped: an estimated duration above 60 seconds is rejected for inline captions, and a caption-stage failure is a soft fail that leaves a clean video_url with captions.status=failed.
What should you test first?
Run one 15-second script through a preview, approve the first frame, and render it at standard. Then compare with plus on the same preview, since stills are reused across tiers. That gives you a real read on speed and quality for your own content, which no changelog can.
Sources
Related posts
More in Comparisons
- HeyGen scene-scoped edits vs fixing one scene in a Sume avatar
HeyGen's Video Agent can edit one scene and keep the rest. Sume has no scene-edit call: structural changes need a new preview; only quality changes late.
- HeyGen's 22 Spanish and 17 Arabic variants vs Sume's language field
HeyGen lists regional variants such as 22 Spanish and 17 Arabic. Sume's TTS takes one BCP-47 language string and does not publish a per-region variant list.
- Higgsfield AI video translator: 18 languages, lip sync, vs Sume
Higgsfield's video translator dubs into 18 languages and re-syncs lips. Sume has no one-call translator; here is what each does, and the Sume steps instead.
- Higgsfield UGC makes 4 generations per run: Sume queue by plan
Higgsfield says a UGC run can produce up to 4 generations. In Sume you submit one job per clip and plan limits set how many run and queue at once.
Written by Sume