HeyGen Video API 5-15 s at 768p vs Sume avatar video 4-60 s

HeyGen's new Video API makes 5-15 second clips at 480p or 768p from a prompt. Sume's avatar video renders scripts of 4-60 seconds at 720p. Which fits your job?

5 min readSume
All posts

HeyGen's new generative Video API makes 5-15 second clips at 480p or 768p from a text prompt, an image or references. Sume's Avatar 1.0 talking video is a different product: a stored avatar speaks a script, the estimated length must land between 4 and 60 seconds, and output is 720p.

So the choice is not about which number is bigger. One tool generates scenes from a prompt; the other makes a specific presenter say specific words.

What does HeyGen's Video API accept?

The changelog entry describes POST /v3/models/videos and GET /v3/models/videos/{video_id} for the heygen-video-1 model, with three modes: text_to_video, image_to_video and reference_to_video. Output is 5 to 15 seconds at 480p or 768p, with prompts up to 5,000 characters. Image-to-video follows the proportions of the source image. Reference-to-video accepts up to nine images, three videos and three audio recordings, twelve references in total. Creation returns HTTP 202 with a video id you poll.

What does Sume's avatar video accept?

POST /v1/avatar-1.0/talking-video takes a ready avatar_handle and exactly one of script or video_inputs. Scripts and multi-scene plans are accepted when Sume estimates the video at 4-60 seconds inclusive. aspect_ratio supports 1:1, 3:4, 9:16, 4:3 and 16:9 (default 9:16), and resolution is currently 720p. Quality is standard, plus (the default) or max.

Optional product_image and scene inputs, either a prompt or a public photo URL, add context. The avatar speaks; there is no free-form prompt describing the whole film.

Two different jobs (HeyGen changelog and Sume docs, read 2026-10-02)
HeyGen Video APISume avatar video
InputPrompt, image, or referencesAvatar handle plus script or video_inputs
Length5-15 seconds4-60 seconds, estimated
Resolution480p or 768p720p
Who is on screenWhatever the prompt generatesOne resolved avatar per video
Returns202 and a video id to poll202 with job id plus status, result and events URLs

When do you pick which?

Pick a prompt-driven generator when you want scenery, motion or a stylised shot with no fixed presenter. Pick the avatar route when the output is a person delivering a script: product pitches, explainers, onboarding, localised greetings.

If you need non-avatar clips on Sume, the Video page covers the video models, which are separate from Avatar 1.0 and have their own limits. Do not expect the avatar route to take a long prompt and invent a scene.

How do you get past 60 seconds on Sume?

You cannot stretch one avatar video beyond the window: longer scripts must be shortened or split into multiple jobs. A common pattern is one job per chapter, then sequencing the finished clips. Each job is its own paid generation with its own reservation, so budget the sum.

Inline captions are also capped: an estimated duration above 60 seconds is rejected for inline captions, and a caption-stage failure is a soft fail that leaves a clean video_url with captions.status=failed.

What should you test first?

Run one 15-second script through a preview, approve the first frame, and render it at standard. Then compare with plus on the same preview, since stills are reused across tiers. That gives you a real read on speed and quality for your own content, which no changelog can.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume