HeyGen liveavatar-gpt-live-demos: where Sume's avatar jobs start

HeyGen's liveavatar-gpt-live-demos repo is a real-time avatar starter. Sume renders avatar video as jobs: submit async, poll or webhook, then fetch the MP4.

4 min readSume
All posts

liveavatar-gpt-live-demos is a starter for a live, spoken conversation with an avatar. Sume's documented avatar surface is different: a render job. You submit a script, poll or take a webhook, then fetch a finished MP4 from media.sume.com. If your product needs the conversation, that part is not on Sume.

The repo facts are from HeyGen's September 2026 release post. Sume facts are from Jobs and results, Webhooks and Generate avatar video, read 2026-10-01.

What does the HeyGen demo contain?

Per the post, on September 10 HeyGen open-sourced a reference where OpenAI's GPT-Live-1, a full-duplex speech-to-speech model, drives a HeyGen LiveAvatar and tool calls appear as animated HyperFrames overlays. Both demo repos are MIT licensed. The post calls them starters, not deployments, and says to add auth before exposing one publicly.

What does a render job need instead?

Nothing in the loop is real time. The three decisions are how you submit, how you learn the result and how you fetch it.

Sume avatar job mechanics, docs read 2026-10-01.
StepWhat the docs say
SubmitPrefer async with an Idempotency-Key for production integrations
Waitwait_timeout_seconds is clamped to 0..30; it bounds the HTTP wait, not the job
Be toldWebhooks send terminal events only (job.completed, job.failed, job.canceled)
FetchGET /v1/jobs/:id/result returns media.sume.com artifacts
LengthScripts and plans must estimate at 4 to 60 seconds

Can a Sume avatar clip play the role of a live reply?

Only as a pre-rendered clip. A render takes a job's worth of time, and there is no progress or partial delivery, so a turn-by-turn conversation needs a different component. A clip fits cases where the reply is known in advance, such as a greeting or product explainer.

Where do I go from here?

For billing contrasts with a voice session, see per-second voice pricing vs per-job audio.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume