Two-host AI video podcast: what Sume avatar videos can do
HeyGen's Video Podcast puts two AI hosts in one studio scene. A Sume avatar video resolves one avatar per final video, so a two-host show is two renders.
A single Sume avatar video has one resolved avatar, so two hosts sharing one frame is not something the docs show. For a two-host format, render each host as its own avatar video and cut between them on a Timeline. HeyGen's July 2026 Video Podcast is a single product feature with two hosts in one scene.
What does HeyGen's Video Podcast do?
HeyGen says it "turns any script, URL, PDF, doc, or topic into a two-host video podcast in minutes", and that "Your two AI hosts share a real studio scene and react to each other". Both quotes are from its July 2026 release notes.
What does a Sume avatar video support?
The avatar video docs say current execution supports "one resolved avatar per final video" and expects scene backgrounds to resolve to one shared scene. Multi-scene video_inputs are ordered scenes of one avatar with voice or silence beats, not two people talking.
| Question | HeyGen (per its notes) | Sume docs |
|---|---|---|
| Two hosts in one video | Yes, in Video Podcast | No: one resolved avatar per final video |
| Shared studio scene | Yes | One shared scene per video |
| Silence beat while the other host talks | Not stated | voice.type "silence" with a duration |
How do I fake a conversation?
Write the dialogue as alternating lines, render each host's lines as separate avatar videos with different avatar_handle values, and use the same scene prompt so the backgrounds match. Then sequence the clips on Timeline 1.0. The result is cuts between hosts, not a two-shot where one reacts while the other speaks; the earlier post on two-speaker avatar videos covers the same idea.
What are the costs of that approach?
Each clip is its own job with its own reservation, and each must be 4-60 estimated seconds. A 10-minute show cut into 30-second turns is 20 requests. Keep the same aspect_ratio across all of them.
Sources
Related posts
More in Sume Avatar 1.0
- Face swap or avatar talking video: which Sume endpoint?
Use face swap when you already have a 4-15 second source video; use the avatar talking video when you have a script. Inputs, limits and what each returns.
- HeyGen 30-minute avatar video: what Sume does in one request
HeyGen says it can make a 30-minute talking video in one pass. One Sume avatar video request covers 4-60 seconds, so longer pieces are split into jobs.
- HeyGen Avatar V's 15-second recording vs a Sume photo avatar
HeyGen's Avatar V starts from a 15-second recording. Sume's avatar API starts from a prompt, traits or one public HTTPS image, with no recording step.
- HeyGen Edit Look: retouching an AI avatar, and the Sume route
HeyGen's Edit Look retouches an existing avatar in place. Sume's docs show no such edit, so the route is a new avatar from a retouched photo and a new handle.
Written by Sume