Interview video with one AI avatar host and guest cards on screen

A final avatar video holds one avatar, so an interview needs a host clip plus guest stills. Use Timeline compose at $0.02 a shot to put a still beside the host.

5 min readSume
All posts

Sume Avatar 1.0 renders one resolved avatar per final video, so an interview is built as separate pieces: host clips from the avatar, and guest material as stills or your own footage. Timeline compose then puts one still and one video on screen together, for a flat $0.02 per shot per the pricing code.

This keeps every piece inside what the docs describe, and it keeps the guest's face real rather than synthetic.

Pieces and where they come from

Each row maps to a documented route, so you can script the whole flow.

Interview parts and the Sume surface for each, read 2026-10-06, from the Avatar 1.0 and Timeline compose docs.
PieceSurfaceLimit
Host question clipsPOST /v1/avatar-1.0/talking-video4-60 s each, one avatar
Guest card (name, quote)A still you import to media.sume.comMust probe as a still
Host beside guest cardPOST /v1/timeline-1.0/compose, operation stack or overlayOne still plus one video; up to 300 s
Final assemblyTimeline 1.0$0.10 per ceil output minute

Flow

Import the guest still first with POST /v1/media-imports, because compose needs both URLs to be this workspace's media.sume.com assets. Compose requires an Idempotency-Key and returns a job you poll at /v1/jobs/:id/status. The operation field selects stack or overlay.

The output length comes from the video layer, so the card stays on screen for the whole clip and cannot make it longer.

Order the work so the cheap steps come first: write the questions, render one host clip, compose it with one card, and only then render the rest.

When this beats a two-avatar render

A separate avatar for the guest means a second job and a second face, and the finished piece is only two talking heads. For a real guest, use their real recording or photo. For a fictional panel, see the two-speaker posts linked below.

If you do need an avatar for both speakers, the cost doubles and the two videos still need a cut between them. That is a different design, covered in the linked posts.

Cost sketch

Four 20-second host questions at standard are about $14.72. Four compose shots add $0.08, and a 3-minute assembly adds $0.30. Treat these as estimates; confirm with a dry run.

Use dry_run in the hosted MCP tools to preview the cost before an agent spends anything, and set max_spend_usd as a cap.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume