Podcast host photo speaks a 12-second highlight: $1.21 on Sume
Cut a 12 second highlight from a podcast video with audio detach ($0.01), then drive a host photo with H3 Max lip sync at 768p ($1.20). Total $1.21, no editor.

The short answer
To turn a podcast highlight into a speaking host photo, detach the audio of that moment from the episode video with POST /v1/audio-detach and a range, then send the WAV and the photo to H3 Max lip sync. A 12-second highlight costs $0.01 for the detach and $1.20 at 768p for the lip sync, so $1.21 in total.
Step 1: cut the audio
Audio detach reads one Sume-hosted video and returns a new audio file, wav by default. The range field takes start and end in seconds. A 12-second highlight starting at 754.2 seconds is { "start": 754.2, "end": 766.2 }. Mono keeps the file small. The source video is untouched.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: highlight-cut" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/episode.mp4",
"range": {"start": 754.2, "end": 766.2},
"channels": "mono"
}'Where the audio lands
The job result has an audio_url, a new artifact on the Sume media host, and duration_seconds. Import the episode first with POST /v1/media-imports if it is not on the media host yet. The detach call is $0.01 a job.
Step 2: lip sync
Send the host photo as image_url and the new audio_url. Pass the real length as duration_seconds: the route needs 5 to 14.8 seconds and rejects anything else instead of clamping it.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: highlight-lipsync" \
-d '{
"image_url": "https://example.com/host.jpg",
"audio_url": "https://media.sume.com/artifacts/artf_demo/highlight.wav",
"duration_seconds": 12,
"resolution": "768p"
}'Choosing the moment
Choose a highlight where the host speaks alone. Crosstalk, laughter and music beds make the mouth movement worse, and the voice of a second person would be lip-synced to the host. If the moment is longer than 14.8 seconds, cut two highlights and join them.
Cost
The rates are per second, rounded up, at list x 1.25.
| Step | Basis | Cost |
|---|---|---|
| Audio detach | per job | $0.01 |
| H3 Max lip sync, 768p | 12 s x $0.10 | $1.20 |
| Total | $1.21 |
Sources
Related posts
More in Use cases
- A podcast intro line in 5 languages: localized sting for 8 cents
Speak a 120-character intro in five languages and concat each ahead of an episode: 600 characters of TTS plus five $0.01 concats is $0.08 on Sume.
- Post-call sales recap video: three-scene AI avatar summary, next step
After a sales call, send a 30-second avatar recap in three scenes: thanks, agreed points, next step. Sume request body, shared-scene rule, cost per prospect.
- Same AI video on TikTok and YouTube Shorts: reach risk?
YouTube's Oct 2026 update targets re-uploads from other creators. Is your own AI video on TikTok and Shorts affected? What is stated, what is not, what to do.
- Pre-render top 20 support answers as avatar clips; agent for the rest
Render the 20 most common answers as Sume avatar clips, serve them by intent, and send the rest to a person. Tavus Griffin is not open to customers.
Written by Sume