Two AI characters in one TikTok scene: one avatar per job, then cut
One Sume avatar job holds one avatar. For a two-person dialogue, make a job per speaker and alternate them in Timeline 1.0. Cost for 60 s of dialogue: $14.82.
Sume Avatar 1.0 supports one resolved avatar per final video, so a two-person dialogue is two avatar jobs cut together in Timeline 1.0. Sixty seconds of dialogue, 30 seconds per speaker, is $14.70 at the plus tier, plus $0.12 for two audio detaches and the join. This works for the talking-head scenes common in micro-dramas, which Metricool reports as gaining traction on TikTok.
Be clear about the limit: the two people never share a frame. It is a shot, reverse shot structure.
The documented limit
The avatar video docs state that only one resolved avatar is supported per final video, and that scene backgrounds must resolve to one shared scene. So each speaker needs their own job, and the two jobs should be set in matching rooms.
Plan the exchange
Write the dialogue as alternating lines, then group them. Each job holds consecutive lines for that speaker only, up to 60 seconds. For shorter exchanges, make more, shorter jobs: each runs 4 to 60 seconds.
| Item | Seconds | Rate | Cost |
|---|---|---|---|
| Speaker A job | 30 | $0.245 per s | $7.35 |
| Speaker B job | 30 | $0.245 per s | $7.35 |
| Two audio detaches | $0.01 each | $0.02 | |
| Timeline join | 60 | $0.10 per min | $0.10 |
| Total | 60 | $14.82 |
Keep the sound in order
Each avatar job returns a video with its own voice. To put them back to back in order, use the video's own audio. The simplest route is to detach the sound from each job with audio detach at $0.01 per job, then pass the files as audio.parts[] (up to 20) in the same order as the video slots. The table above includes them.
{
"audio": {"duration_seconds": 60, "parts": [
{"url": "https://media.sume.com/artifacts/artf_a/a-line1.wav"},
{"url": "https://media.sume.com/artifacts/artf_b/b-line1.wav"}]},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_a/a-line1.mp4", "start": 0, "duration": 30},
{"source_url": "https://media.sume.com/artifacts/artf_b/b-line1.mp4", "start": 30, "duration": 30}]
}Where this falls short
Quick back and forth needs many short jobs, and each job has a 4 second minimum. Overlapping speech is not possible this way. For two people in one frame use a video model with image references, described in multiple characters in one AI video.
Sources
Related posts
More in Sume Avatar 1.0
- YouTube avatar vs Sume Avatar: selfie capture or prompt and photo
YouTube's avatar is made once from your own face and voice and used in its AI tools. How it is created, its limits, and how Sume's Avatar 1.0 differs.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
Written by Sume