HeyGen liveavatar-gpt-live-demos: where Sume's avatar jobs start
HeyGen's liveavatar-gpt-live-demos repo is a real-time avatar starter. Sume renders avatar video as jobs: submit async, poll or webhook, then fetch the MP4.
liveavatar-gpt-live-demos is a starter for a live, spoken conversation with an avatar. Sume's documented avatar surface is different: a render job. You submit a script, poll or take a webhook, then fetch a finished MP4 from media.sume.com. If your product needs the conversation, that part is not on Sume.
The repo facts are from HeyGen's September 2026 release post. Sume facts are from Jobs and results, Webhooks and Generate avatar video, read 2026-10-01.
What does the HeyGen demo contain?
Per the post, on September 10 HeyGen open-sourced a reference where OpenAI's GPT-Live-1, a full-duplex speech-to-speech model, drives a HeyGen LiveAvatar and tool calls appear as animated HyperFrames overlays. Both demo repos are MIT licensed. The post calls them starters, not deployments, and says to add auth before exposing one publicly.
What does a render job need instead?
Nothing in the loop is real time. The three decisions are how you submit, how you learn the result and how you fetch it.
| Step | What the docs say |
|---|---|
| Submit | Prefer async with an Idempotency-Key for production integrations |
| Wait | wait_timeout_seconds is clamped to 0..30; it bounds the HTTP wait, not the job |
| Be told | Webhooks send terminal events only (job.completed, job.failed, job.canceled) |
| Fetch | GET /v1/jobs/:id/result returns media.sume.com artifacts |
| Length | Scripts and plans must estimate at 4 to 60 seconds |
Can a Sume avatar clip play the role of a live reply?
Only as a pre-rendered clip. A render takes a job's worth of time, and there is no progress or partial delivery, so a turn-by-turn conversation needs a different component. A clip fits cases where the reply is known in advance, such as a greeting or product explainer.
Where do I go from here?
For billing contrasts with a voice session, see per-second voice pricing vs per-job audio.
Sources
Related posts
More in Developers
- HeyGen create avatar from a text prompt API vs Sume
HeyGen POST /v3/avatars type prompt takes up to 1000 characters and an aspect_ratio. Sume creates a text-only avatar with a Prompt input on avatar-1.0/generate.
- HeyGen Stripe Projects API key vs how you get a Sume key
HeyGen lets an agent provision a key with stripe projects add heygen/api. Sume keys are created in the dashboard, workspace-scoped and shown once.
- HeyGen create template via API: Sume has video_inputs, not templates
HeyGen now creates templates over the API and binds variables by element_id. Sume has no stored template object; send ordered video_inputs on each request.
- HeyGen studio video scene limit: 50 scenes, and Sume's 4-60 s plan
HeyGen studio videos allow 50 scenes and 30 minutes per scene. Sume multi-scene video_inputs share one 4-60 second window and one resolved avatar.
Written by Sume