Lip sync AI music video: split a song into 5 to 14.8 second windows
MiniMax H3 Max lip sync accepts 5 to 14.8 seconds of audio per clip. Split a 60-second song into five 12-second ranges: about $6.24 at 768p on Sume.

"Lip sync AI music video" is another phrase from today's autocomplete (read 2026-10-07). The constraint that shapes the answer on Sume is a hard one: MiniMax H3 Max lip sync takes audio of 5 to 14.8 seconds, because the provider rejects shorter audio and silently clips anything longer. Sume refuses both edges at admission instead of clamping. A song does not fit in one call, so you cut it into windows and make one clip per window.
The cutting is already a Sume job. Timeline audio with operation: "split" takes one Sume-hosted file and up to 20 ranges, each {start, end}, and returns one durable URL per range.
Five windows for a 60-second song
Choose windows that fall on a phrase end, not on a fixed count. For the cost example, five equal 12-second ranges: 0-12, 12-24, 24-36, 36-48, 48-60. Every range is above 5 and below 14.8 seconds. Keep the WAV format (the default) for the split, since the docs say a WAV that drives lip sync should stay sample-exact.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: song-split-001" \
-d '{"operation": "split", "url": "https://media.sume.com/artifacts/artf_demo/song.wav", "ranges": [{"start": 0, "end": 12}, {"start": 12, "end": 24}, {"start": 24, "end": 36}, {"start": 36, "end": 48}, {"start": 48, "end": 60}]}'What it costs
| Step | Basis | 480p | 768p | 1080p |
|---|---|---|---|---|
| Music Router track (skip if you own it) | $0.125 per generation | $0.125 | $0.125 | $0.125 |
| Split into five ranges | $0.01 per job | $0.01 | $0.01 | $0.01 |
| Five lip-sync clips, 60 audio seconds | $0.0625, $0.10, $0.20 per second | $3.75 | $6.00 | $12.00 |
| Timeline render, 60 s | $0.10 per ceil(output minute) | $0.10 | $0.10 | $0.10 |
| Total | $3.985 | $6.235 | $12.235 |
Practical limits
- Each lip-sync request also needs a face: an image, and audio that is Sume-hosted and under 10 MB.
- If you generate the song on Sume, leave the "no vocals" clause out of the prompt; the docs recommend ending an instrumental brief with it, so omitting it is how you ask for voice. Check that the sung words are what you want before you spend on clips.
- Billing counts whole audio seconds, so a 12.3-second range is billed as 13.
Sources
Related posts
More in Use cases
- Listing walkthrough from 8 photos: seven 4-second Wan 3.0 transitions
Eight room photos become seven Wan 3.0 first-to-last-frame transitions of 4 seconds. $3.50 at 720p plus $0.10 for the timeline render on Sume.
- Live stream avatar: queue pre-rendered reaction clips instead
For a stream overlay, render ten short Sume avatar clips ahead of time and trigger them from chat events. Ten 6-second clips cost $14.70 on plus.
- Lock the look with one image, then animate it on three video models
Make one approved image, then use it as first_frame on wan-3.0, kling-3 and minimax-h3-max. Same start, three motions, one prompt, so you can compare.
- Lofi rain-on-window loop with sound: Gemini Omni 10 seconds
Make a 10-second lofi rainy window clip with Gemini Omni on Sume. Ambient audio wording, one image as first and last frame for a loop, and the cost to draft.
Written by Sume