HeyGen Video Agent prompt-to-video vs a Sume script-driven avatar job

HeyGen's v3 quick start creates a video from a prompt and returns a session_id. Sume's avatar route takes a script and avatar handle and returns a job.

5 min readSume
All posts

HeyGen's quick start calls POST /v3/video-agents with a prompt, labels it the flagship path (prompt in, finished video out), and returns a session_id you poll until a video_id appears. Sume works the other way round: you pick a ready avatar, write the exact script or a list of scenes, and POST /v1/avatar-1.0/talking-video returns a job you poll or receive by webhook. If you want an agent to decide the wording, HeyGen's flow fits; if the words are legal-approved and must be spoken as written, a script-first call is safer.

HeyGen source: API quick start, read 2026-10-04. Sume sources: Create new avatar and Generate avatar video.

Two request shapes

The quick start describes a two-stage read: poll the session until a video id shows up, then poll the video until it is completed or failed. Sume has one handle, the job, and three reads on it: status, events and result.

Prompt-first vs script-first (read 2026-10-04)
ItemHeyGen v3 Video AgentSume avatar talking-video
Main inputA promptscript, or ordered video_inputs
Who writes the wordsThe agent, from the promptYou, exactly
Handle returnedsession_id, then video_idJob id, plus avatar_video_id resource
Auth headerX-Api-KeyAuthorization: Bearer
CompletionPoll, or callback_urlPoll, or mode: webhook with webhook_url

When the exact words matter

Regulated claims, price statements and disclaimers should not be paraphrased. With Sume you supply the script, and for multi-scene work you can mark a scene as a silent beat with voice.type: "silence" and a required duration. The estimated length must land between 4 and 60 seconds, so a long script has to be split before you submit.

  • Write and approve the script outside the model call.
  • Put the legal line in its own scene so it cannot be trimmed.
  • Use a preview still to check the frame before the full render.

When a prompt is the better input

For exploratory drafts, a prompt that lets an agent propose structure saves a writer's time. Sume reaches the same place through its Agents surface and the hosted tools, but the plain avatar route stays script-driven. Use the Format API when you want a saved recipe that turns a product URL into a video, and keep the avatar route for a controlled presenter clip.

Bottom line

Choose by who owns the words. If the model should write them, test HeyGen's v3 flow. If your team writes them, use Sume's script route and keep the approval step before the render.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume