Two-speaker dialogue video from two photos: $1.80 on Sume
Lip-sync each speaker's photo with H3 Max, then join the two clips on a Timeline render. An 8 s and a 9 s line at 768p cost $1.70, plus $0.10 to join.

The short answer
To make a two-speaker dialogue from two photos, lip-sync each speaker's line with H3 Max, then join the clips with a Timeline render. An 8-second and a 9-second line at 768p cost $1.70 (17 seconds at $0.10), and the render adds $0.10 for one output minute. Total: $1.80. Each audio line must run 5 to 14.8 seconds.
Step 1: one lip-sync job per line
Record or generate each speaker's line as its own audio file, 5 to 14.8 seconds each, and import both to the Sume media host. Then submit one lip-sync job per speaker with that speaker's photo. Each job is billed on its rounded-up seconds: 8 s and 9 s at $0.10 a second.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dialogue-a" \
-d '{
"image_url": "https://example.com/speaker-a.jpg",
"audio_url": "https://media.sume.com/artifacts/artf_demo/line-a.wav",
"duration_seconds": 8,
"resolution": "768p"
}'Step 2: the second speaker
Repeat with the second photo, the second audio and "duration_seconds": 9. Keep both photos at the same framing and resolution so the cut between speakers does not jump. Each clip takes the image's aspect ratio, which must be between 0.4 and 2.5.
Step 3: join them
The join uses the Timeline 1.0 render. The audio spine is the two original lines as gapless audio.parts[], and the two lip-sync videos are video[] slots whose starts match the line lengths. video[0].start must be 0, and each later start must be larger.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dialogue-join" \
-d '{
"audio": {"duration_seconds": 17, "parts": [
{"url": "https://media.sume.com/artifacts/artf_demo/line-a.wav", "duration": 8},
{"url": "https://media.sume.com/artifacts/artf_demo/line-b.wav", "duration": 9}]},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/clip-a.mp4", "start": 0, "duration": 8},
{"source_url": "https://media.sume.com/artifacts/artf_demo/clip-b.mp4", "start": 8, "duration": 9}]
}'What it costs
Timeline billing is $0.10 per output minute, rounded up, so 17 seconds is one minute. The default output is 1080x1920, so a landscape dialogue needs output.width and output.height. Add fit: "contain" or "blur" on a slot if the clips would otherwise crop.
| Step | Basis | Cost |
|---|---|---|
| Lip sync, line A | 8 s x $0.10 | $0.80 |
| Lip sync, line B | 9 s x $0.10 | $0.90 |
| Timeline render | 1 output minute x $0.10 | $0.10 |
| Total | $1.80 |
Longer conversations
For a longer scene, add more lines and more slots, up to 20 audio parts and 200 video slots in one render. See two-speaker conversation video for the avatar route.
Sources
Related posts
More in Use cases
- UGC-style Black Friday ad in three parts: 4, 8 and 3 seconds for $2.18
TikTok's Black Friday guide wants a 3-5 s opener, a benefits middle and a direct close. Three Omni 720p clips (4+8+3 s), a stitch and captions cost $2.18.
- UK drip pricing: put the delivery fee in the Black Friday video price
The CMA wants the total price, mandatory fees included, from early advertising on. Build one video price card that shows it. Read 2026-10-05.
- Unboxing video music that stays under your voice: duck_db settings
Pick duck_db and gain_db on Sume Timeline so a $0.125 AI music bed dips under speech and returns between lines, plus the duck error to expect.
- University Open House Welcome With an AI Avatar: Script and Cost
A university open-house welcome from an AI avatar: a scripted 3-scene plan, request JSON, Sume cost per quality tier, and how it differs from live avatars.
Written by Sume