Lip sync AI music video: split a song into 5 to 14.8 second windows

MiniMax H3 Max lip sync accepts 5 to 14.8 seconds of audio per clip. Split a 60-second song into five 12-second ranges: about $6.24 at 768p on Sume.

6 min readSume
All posts

"Lip sync AI music video" is another phrase from today's autocomplete (read 2026-10-07). The constraint that shapes the answer on Sume is a hard one: MiniMax H3 Max lip sync takes audio of 5 to 14.8 seconds, because the provider rejects shorter audio and silently clips anything longer. Sume refuses both edges at admission instead of clamping. A song does not fit in one call, so you cut it into windows and make one clip per window.

The cutting is already a Sume job. Timeline audio with operation: "split" takes one Sume-hosted file and up to 20 ranges, each {start, end}, and returns one durable URL per range.

Five windows for a 60-second song

Choose windows that fall on a phrase end, not on a fixed count. For the cost example, five equal 12-second ranges: 0-12, 12-24, 24-36, 36-48, 48-60. Every range is above 5 and below 14.8 seconds. Keep the WAV format (the default) for the split, since the docs say a WAV that drives lip sync should stay sample-exact.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: song-split-001" \
  -d '{"operation": "split", "url": "https://media.sume.com/artifacts/artf_demo/song.wav", "ranges": [{"start": 0, "end": 12}, {"start": 12, "end": 24}, {"start": 24, "end": 36}, {"start": 36, "end": 48}, {"start": 48, "end": 60}]}'

What it costs

60-second performance video, Sume catalog read 2026-10-07
StepBasis480p768p1080p
Music Router track (skip if you own it)$0.125 per generation$0.125$0.125$0.125
Split into five ranges$0.01 per job$0.01$0.01$0.01
Five lip-sync clips, 60 audio seconds$0.0625, $0.10, $0.20 per second$3.75$6.00$12.00
Timeline render, 60 s$0.10 per ceil(output minute)$0.10$0.10$0.10
Total$3.985$6.235$12.235

Practical limits

  • Each lip-sync request also needs a face: an image, and audio that is Sume-hosted and under 10 MB.
  • If you generate the song on Sume, leave the "no vocals" clause out of the prompt; the docs recommend ending an instrumental brief with it, so omitting it is how you ask for voice. Check that the sung words are what you want before you spend on clips.
  • Billing counts whole audio seconds, so a 12.3-second range is billed as 13.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume