Two AI characters in one TikTok scene: one avatar per job, then cut

One Sume avatar job holds one avatar. For a two-person dialogue, make a job per speaker and alternate them in Timeline 1.0. Cost for 60 s of dialogue: $14.82.

5 min readSume
All posts

Sume Avatar 1.0 supports one resolved avatar per final video, so a two-person dialogue is two avatar jobs cut together in Timeline 1.0. Sixty seconds of dialogue, 30 seconds per speaker, is $14.70 at the plus tier, plus $0.12 for two audio detaches and the join. This works for the talking-head scenes common in micro-dramas, which Metricool reports as gaining traction on TikTok.

Be clear about the limit: the two people never share a frame. It is a shot, reverse shot structure.

The documented limit

The avatar video docs state that only one resolved avatar is supported per final video, and that scene backgrounds must resolve to one shared scene. So each speaker needs their own job, and the two jobs should be set in matching rooms.

Plan the exchange

Write the dialogue as alternating lines, then group them. Each job holds consecutive lines for that speaker only, up to 60 seconds. For shorter exchanges, make more, shorter jobs: each runs 4 to 60 seconds.

Dialogue cost, 60 s total, plus tier (as of 2026-10-03)
ItemSecondsRateCost
Speaker A job30$0.245 per s$7.35
Speaker B job30$0.245 per s$7.35
Two audio detaches$0.01 each$0.02
Timeline join60$0.10 per min$0.10
Total60$14.82

Keep the sound in order

Each avatar job returns a video with its own voice. To put them back to back in order, use the video's own audio. The simplest route is to detach the sound from each job with audio detach at $0.01 per job, then pass the files as audio.parts[] (up to 20) in the same order as the video slots. The table above includes them.

{
  "audio": {"duration_seconds": 60, "parts": [
    {"url": "https://media.sume.com/artifacts/artf_a/a-line1.wav"},
    {"url": "https://media.sume.com/artifacts/artf_b/b-line1.wav"}]},
  "video": [
    {"source_url": "https://media.sume.com/artifacts/artf_a/a-line1.mp4", "start": 0, "duration": 30},
    {"source_url": "https://media.sume.com/artifacts/artf_b/b-line1.mp4", "start": 30, "duration": 30}]
}

Where this falls short

Quick back and forth needs many short jobs, and each job has a 4 second minimum. Overlapping speech is not possible this way. For two people in one frame use a video model with image references, described in multiple characters in one AI video.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume