Detach 13 seconds of speech from a clip and lip-sync it on H3 Max

Sume can detach audio from a hosted clip by time range, then drive a talking face with H3 Max lip-sync. $1.31 for 13 seconds; rights and limits explained.

5 min readSume
All posts

Yes: audio detach cuts a time range out of a Sume-hosted video as a wav, and H3 Max lip-sync can then animate a still to that audio, as long as the piece is 5 to 14.8 seconds. A 13 second piece costs $0.01 for the detach and $1.30 for the lip-sync at 768p, so $1.31 in total.

Use this only with speech you have the right to reuse and a face you have the right to animate. The limits below come from the Audio detach page and the Models overview, checked 2026-10-10.

Step 1: detach the range

POST /v1/audio-detach takes a video_url on the Sume media host, plus an optional range with start and end. The default output is sample-exact wav (pcm_s16le); mp3 at 128 kbps is the alternative. The source may be up to 1,800 seconds and the output up to 900. Errors you may meet include detach_source_has_no_audio and audio_detach_range_empty.

  • Set range.start and range.end so the piece is between 5 and 14.8 seconds.
  • Keep the wav default; it is sample-exact.
  • Use channels: "mono" if you want a smaller file.

Step 2: size check

Uncompressed 16-bit wav is large but well inside the cap: mono at 44.1 kHz is about 88 KB per second, so 14.8 seconds is about 1.3 MB. Stereo at 48 kHz is about 192 KB per second, or about 2.8 MB for 14.8 seconds. The Fabric route caps audio at 10 MB; both stay under it, though the H3 Max route's own body should be checked in the schema.

Approximate wav sizes for 14.8 s of 16-bit PCM (arithmetic)
FormatBytes per secondSize for 14.8 s
Mono 44.1 kHz88,200about 1.3 MB
Stereo 48 kHz192,000about 2.8 MB
Mono 16 kHz32,000about 0.5 MB

Step 3: lip-sync and cost

Send the detached audio and a still to POST /v1/minimax/h3-max/lip-sync. H3 Max bills per ceiling audio second: 480p $0.0625, 768p $0.10, 1080p $0.20. Image aspect must be between 0.4 and 2.5. A 13.0 second piece rounds to 13 seconds; 13.2 would bill 14.

Cost for a 13 s range (Sume rates, checked 2026-10-10)
ItemMathCost
Audio detachper job$0.01
H3 Max, 480p13 x $0.0625$0.8125
H3 Max, 768p13 x $0.10$1.30
H3 Max, 1080p13 x $0.20$2.60
Total at 768p$1.31

Rights and honesty

Speech and faces belong to people. Detach only audio you own or have permission to reuse, animate only a face with consent, and label the result where a platform expects it. If the aim is a translated or re-voiced version, the dubbing route (detach, transcribe, translate, TTS) is the better fit, because it replaces the voice instead of reusing it.

When the clip is a poor fit

Detached speech carries its room and its background noise. If the source has music under the voice, the detached wav has music under the voice too, and H3 Max will animate to the whole signal. A quiet interview or a dry voice-over works best; a noisy event clip is a poor source. Preview the wav before you pay for the lip-sync.

Also check the length of the piece against the 5 second minimum. A 3 second reaction line is below the floor and the request will not pass; pad with a real pause from the source rather than silence you create.

  • Prefer dry speech without music underneath.
  • Preview the detached wav first.
  • Keep each piece between 5 and 14.8 seconds.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume