Interview video with one AI avatar host and guest cards on screen
A final avatar video holds one avatar, so an interview needs a host clip plus guest stills. Use Timeline compose at $0.02 a shot to put a still beside the host.
Sume Avatar 1.0 renders one resolved avatar per final video, so an interview is built as separate pieces: host clips from the avatar, and guest material as stills or your own footage. Timeline compose then puts one still and one video on screen together, for a flat $0.02 per shot per the pricing code.
This keeps every piece inside what the docs describe, and it keeps the guest's face real rather than synthetic.
Pieces and where they come from
Each row maps to a documented route, so you can script the whole flow.
| Piece | Surface | Limit |
|---|---|---|
| Host question clips | POST /v1/avatar-1.0/talking-video | 4-60 s each, one avatar |
| Guest card (name, quote) | A still you import to media.sume.com | Must probe as a still |
| Host beside guest card | POST /v1/timeline-1.0/compose, operation stack or overlay | One still plus one video; up to 300 s |
| Final assembly | Timeline 1.0 | $0.10 per ceil output minute |
Flow
Import the guest still first with POST /v1/media-imports, because compose needs both URLs to be this workspace's media.sume.com assets. Compose requires an Idempotency-Key and returns a job you poll at /v1/jobs/:id/status. The operation field selects stack or overlay.
The output length comes from the video layer, so the card stays on screen for the whole clip and cannot make it longer.
Order the work so the cheap steps come first: write the questions, render one host clip, compose it with one card, and only then render the rest.
When this beats a two-avatar render
A separate avatar for the guest means a second job and a second face, and the finished piece is only two talking heads. For a real guest, use their real recording or photo. For a fictional panel, see the two-speaker posts linked below.
If you do need an avatar for both speakers, the cost doubles and the two videos still need a cut between them. That is a different design, covered in the linked posts.
Cost sketch
Four 20-second host questions at standard are about $14.72. Four compose shots add $0.08, and a 3-minute assembly adds $0.30. Treat these as estimates; confirm with a dry run.
Use dry_run in the hosted MCP tools to preview the cost before an agent spends anything, and set max_spend_usd as a cap.
Sources
Related posts
More in Sume Avatar 1.0
- Pipecat avatar TTFB metrics vs timing a Sume avatar job
Pipecat's Tavus service reports TTFB from TTSStartedFrame and BotStartedSpeakingFrame. A Sume avatar job has no first byte to time: measure submit to completed.
- Pipecat HeyGen LiveAvatar: duration by tier vs a 60 second Sume job
In Pipecat, HeyGen LiveAvatar session length depends on your subscription tier. On Sume the length rule is fixed: 4 to 60 seconds per job, priced per second.
- Pipecat Simli max_session_length vs Sume's 4 to 60 second clip
Simli in Pipecat caps a live session with max_session_length and max_idle_time. Sume has no session: an avatar job is a 4 to 60 second file.
- Pipecat TavusVideoService live avatar, or a recorded Sume clip?
Pipecat's Tavus service speaks your agent's TTS live over WebRTC. A Sume avatar job is a 4 to 60 second recorded MP4. How to pick, by what the viewer does.
Written by Sume