Podcast host photo speaks a 12-second highlight: $1.21 on Sume

Cut a 12 second highlight from a podcast video with audio detach ($0.01), then drive a host photo with H3 Max lip sync at 768p ($1.20). Total $1.21, no editor.

5 min readSume
All posts

The short answer

To turn a podcast highlight into a speaking host photo, detach the audio of that moment from the episode video with POST /v1/audio-detach and a range, then send the WAV and the photo to H3 Max lip sync. A 12-second highlight costs $0.01 for the detach and $1.20 at 768p for the lip sync, so $1.21 in total.

Step 1: cut the audio

Audio detach reads one Sume-hosted video and returns a new audio file, wav by default. The range field takes start and end in seconds. A 12-second highlight starting at 754.2 seconds is { "start": 754.2, "end": 766.2 }. Mono keeps the file small. The source video is untouched.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: highlight-cut" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/episode.mp4",
    "range": {"start": 754.2, "end": 766.2},
    "channels": "mono"
  }'

Where the audio lands

The job result has an audio_url, a new artifact on the Sume media host, and duration_seconds. Import the episode first with POST /v1/media-imports if it is not on the media host yet. The detach call is $0.01 a job.

Step 2: lip sync

Send the host photo as image_url and the new audio_url. Pass the real length as duration_seconds: the route needs 5 to 14.8 seconds and rejects anything else instead of clamping it.

curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: highlight-lipsync" \
  -d '{
    "image_url": "https://example.com/host.jpg",
    "audio_url": "https://media.sume.com/artifacts/artf_demo/highlight.wav",
    "duration_seconds": 12,
    "resolution": "768p"
  }'

Choosing the moment

Choose a highlight where the host speaks alone. Crosstalk, laughter and music beds make the mouth movement worse, and the voice of a second person would be lip-synced to the host. If the moment is longer than 14.8 seconds, cut two highlights and join them.

Cost

The rates are per second, rounded up, at list x 1.25.

Read from the Sume docs on 2026-10-05
StepBasisCost
Audio detachper job$0.01
H3 Max lip sync, 768p12 s x $0.10$1.20
Total$1.21

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume