Two-speaker dialogue video from two photos: $1.80 on Sume

Lip-sync each speaker's photo with H3 Max, then join the two clips on a Timeline render. An 8 s and a 9 s line at 768p cost $1.70, plus $0.10 to join.

5 min readSume
All posts

The short answer

To make a two-speaker dialogue from two photos, lip-sync each speaker's line with H3 Max, then join the clips with a Timeline render. An 8-second and a 9-second line at 768p cost $1.70 (17 seconds at $0.10), and the render adds $0.10 for one output minute. Total: $1.80. Each audio line must run 5 to 14.8 seconds.

Step 1: one lip-sync job per line

Record or generate each speaker's line as its own audio file, 5 to 14.8 seconds each, and import both to the Sume media host. Then submit one lip-sync job per speaker with that speaker's photo. Each job is billed on its rounded-up seconds: 8 s and 9 s at $0.10 a second.

curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dialogue-a" \
  -d '{
    "image_url": "https://example.com/speaker-a.jpg",
    "audio_url": "https://media.sume.com/artifacts/artf_demo/line-a.wav",
    "duration_seconds": 8,
    "resolution": "768p"
  }'

Step 2: the second speaker

Repeat with the second photo, the second audio and "duration_seconds": 9. Keep both photos at the same framing and resolution so the cut between speakers does not jump. Each clip takes the image's aspect ratio, which must be between 0.4 and 2.5.

Step 3: join them

The join uses the Timeline 1.0 render. The audio spine is the two original lines as gapless audio.parts[], and the two lip-sync videos are video[] slots whose starts match the line lengths. video[0].start must be 0, and each later start must be larger.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: dialogue-join" \
  -d '{
    "audio": {"duration_seconds": 17, "parts": [
      {"url": "https://media.sume.com/artifacts/artf_demo/line-a.wav", "duration": 8},
      {"url": "https://media.sume.com/artifacts/artf_demo/line-b.wav", "duration": 9}]},
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/clip-a.mp4", "start": 0, "duration": 8},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/clip-b.mp4", "start": 8, "duration": 9}]
  }'

What it costs

Timeline billing is $0.10 per output minute, rounded up, so 17 seconds is one minute. The default output is 1080x1920, so a landscape dialogue needs output.width and output.height. Add fit: "contain" or "blur" on a slot if the clips would otherwise crop.

Billed prices read from the Sume docs on 2026-10-05
StepBasisCost
Lip sync, line A8 s x $0.10$0.80
Lip sync, line B9 s x $0.10$0.90
Timeline render1 output minute x $0.10$0.10
Total$1.80

Longer conversations

For a longer scene, add more lines and more slots, up to 20 audio parts and 200 video slots in one render. See two-speaker conversation video for the avatar route.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume