Two-person micro-drama dialogue on Avatar 1.0: one avatar per video

Avatar 1.0 renders one avatar per final video, so a two-character scene is two jobs. A 30-second exchange costs $7.35 at plus and $16.50 at max.

5 min readSume
All posts

Can Avatar 1.0 put two characters in one micro-drama scene? Not in one video: the current execution supports one resolved avatar for each final video, so a two-person exchange is two jobs (one per speaker) that you cut together afterward. Two 15-second jobs at the plus tier cost $7.35 in total, and the same 30 seconds at max costs $16.50.

Micro-drama is built from short dialogue exchanges, and dialogue is the expensive part. It helps to know how the Avatar 1.0 contract splits a conversation before you write the script.

What the contract says

The talking-video route is POST /v1/avatar-1.0/talking-video. You give it a top-level avatar_handle and either a script or ordered video_inputs. The docs state the limit directly: the current execution supports one resolved avatar for each final video, and it expects scene backgrounds to resolve to one shared scene.

That is a design constraint, not a bug to work around. Treat each speaker as a separate render with its own handle, then interleave the clips in the edit. Each job has its own Idempotency-Key, so a retry of speaker B never re-bills speaker A.

  • One avatar_handle per job; the same handle can be reused across jobs.
  • Total planned duration must land in the 4-60 second window.
  • Use voice type silence for a beat with no speech inside a single job; use a second job for a different speaker.

Cost of a 30-second exchange

Rates come from the Sume catalog: $0.184, $0.245 and $0.55 per second at standard, plus and max with no product image. A 30-second exchange split evenly between two characters is 30 billable seconds in total, because the jobs are priced per second of output.

Uneven splits cost the same in total. If speaker A holds 20 seconds and speaker B 10, the bill is still 30 seconds at your chosen tier. The only way to change the total is to cut lines or drop a tier.

Two-speaker, 30-second scene, no product image (rates read 2026-10-07)
TierRate per secondSpeaker A 15 sSpeaker B 15 sTotal
standard$0.184$2.76$2.76$5.52
plus (default)$0.245$3.675$3.675$7.35
max$0.55$8.25$8.25$16.50

Assembling the cut

Timeline 1.0 takes one audio spine plus ordered video slots and returns one MP4. Its render rate is $0.10 per ceil output minute, and POST /v1/timeline-1.0/plan is an unbilled compile preflight, so you can check the sequence before paying. Every URL you send must already be a media.sume.com artifact or asset of your workspace.

Because the spine is one audio file, plan the dialogue as a voice track first and let each avatar slot cover its own lines. Stable refusal codes such as segment_overlap and timeline_must_start_at_zero make the plan call a cheap way to catch a bad cut list.

Draft cheap, finish once

Run the dialogue at standard while you tune pacing, then re-render the lines that survive at plus or max. You can approve first-frame preview stills before the full render, and the final quality tier you pass to generate-video can differ from the preview's, so you do not rebuild the shot list when you move up a tier.

A lead who walks through places rather than talking to camera is a video-model job, not an avatar job.

Budget for retakes

Plan for one retake per speaker. At plus, a 15-second retake is $3.675, so a 30-second exchange with one retake each costs $14.70 instead of $7.35. Approving the first frame in a preview before rendering lowers the chance of a retake.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume