HeyGen Video Agent prompt-to-video vs a Sume script-driven avatar job
HeyGen's v3 quick start creates a video from a prompt and returns a session_id. Sume's avatar route takes a script and avatar handle and returns a job.
HeyGen's quick start calls POST /v3/video-agents with a prompt, labels it the flagship path (prompt in, finished video out), and returns a session_id you poll until a video_id appears. Sume works the other way round: you pick a ready avatar, write the exact script or a list of scenes, and POST /v1/avatar-1.0/talking-video returns a job you poll or receive by webhook. If you want an agent to decide the wording, HeyGen's flow fits; if the words are legal-approved and must be spoken as written, a script-first call is safer.
HeyGen source: API quick start, read 2026-10-04. Sume sources: Create new avatar and Generate avatar video.
Two request shapes
The quick start describes a two-stage read: poll the session until a video id shows up, then poll the video until it is completed or failed. Sume has one handle, the job, and three reads on it: status, events and result.
| Item | HeyGen v3 Video Agent | Sume avatar talking-video |
|---|---|---|
| Main input | A prompt | script, or ordered video_inputs |
| Who writes the words | The agent, from the prompt | You, exactly |
| Handle returned | session_id, then video_id | Job id, plus avatar_video_id resource |
| Auth header | X-Api-Key | Authorization: Bearer |
| Completion | Poll, or callback_url | Poll, or mode: webhook with webhook_url |
When the exact words matter
Regulated claims, price statements and disclaimers should not be paraphrased. With Sume you supply the script, and for multi-scene work you can mark a scene as a silent beat with voice.type: "silence" and a required duration. The estimated length must land between 4 and 60 seconds, so a long script has to be split before you submit.
- Write and approve the script outside the model call.
- Put the legal line in its own scene so it cannot be trimmed.
- Use a preview still to check the frame before the full render.
When a prompt is the better input
For exploratory drafts, a prompt that lets an agent propose structure saves a writer's time. Sume reaches the same place through its Agents surface and the hosted tools, but the plain avatar route stays script-driven. Use the Format API when you want a saved recipe that turns a product URL into a video, and keep the avatar route for a controlled presenter clip.
Bottom line
Choose by who owns the words. If the model should write them, test HeyGen's v3 flow. If your team writes them, use Sume's script route and keep the approval step before the render.
Sources
Related posts
More in Comparisons
- Ideogram Plus or Pro vs API per image: what the page shows
Ideogram sells Plus at $15 and Pro at $42 a month in credits, and its page lists no per-image API price. How to compare it with a per-image row, with a script.
- Is there a Runway Agent API? What the help pages say, and Sume's
The Runway pages I read describe Workflow Endpoints and an MCP, not an Agent API. Sume's Agent Completions takes a required spend cap and returns a run.
- Kling 4.0 voice reference vs Sume reference_audio_urls
Kling 4.0 Omni Reference accepts voice references. Sume takes 1-3 reference_audio_urls on models such as MiniMax H3; here is what each source documents.
- Kling 4.0 vs Kling 3.0: the spec differences in one dated table
Length, resolution, HDR, references, keyframes, audio and prompt size, Kling 4.0 against 3.0 as stated on Kling's pages, plus how each maps to a Sume request.
Written by Sume