Lip-sync a 20-second line: H3 Max takes 14.8 s of audio, so split it

MiniMax H3 Max lip sync on Sume accepts 5 to 14.8 seconds of audio. How to cut a 20-second line into two takes, or use Fabric for a longer face-to-camera line.

5 min readSume
All posts

The limit and the workaround

Sume's POST /v1/minimax/h3-max/lip-sync takes audio of 5 to 14.8 seconds, so a 20-second line must be cut into two takes of at least 5 seconds each, and each take run as its own job. Cut at a sentence boundary so nobody sees a word split in half, then join the two videos.

If the line is long and the speaker is a still photo, the other route is POST /v1/veed/fabric-1.0, which takes a still and audio of up to 10 MB with no 14.8-second ceiling.

Pick the route

The right route depends on what you start with: a video of a person, or a still.

Lip-sync routes on Sume by input and audio length (read 2026-10-04)
RouteStarts fromAudio ruleNotes
H3 Max lip syncA video5 to 14.8 seconds480p, 768p or 1080p
VEED Fabric 1.0An accepted stillSume media host audio, up to 10 MBStill plus audio
Avatar talking videoA Sume avatar4 to 60 secondsAvatar plus script

Cutting a 20-second line in two

Write the script in two parts before you generate the voice, so each part is a whole thought of roughly 7 to 12 seconds. Generate each part as its own tts-1.0 job. This is better than cutting finished audio, because the voice finishes a sentence naturally. If you already have one 20-second file, use POST /v1/timeline-1.0/audio with operation: "split" to slice it into ranges at $0.01 flat, and choose a silent point between sentences.

Send each audio part with the same video to its own lip-sync job. Make sure each part is within 5 to 14.8 seconds, because a short tail under 5 seconds is refused. Then join the two clips in order with Sume's timeline. Expect a slight change of expression at the join; hide it with a cut-away or a caption. If the two takes must look continuous, start the second take from a frame close to where the first ended, and keep lighting, framing and resolution identical so the seam is hard to spot. Cost scales with the number of takes, so a 40-second line would need three or four, and at that length Fabric is simpler.

A rule from the docs

Sume's docs say that a face that talks is Fabric with an accepted still plus TTS, and that video models do not lip-sync to TTS. So do not expect a text-to-video model to sync a voice file you provide; use a lip-sync route for that.

  • Check the duration_seconds of each audio part before you submit.
  • Pay attention to price: H3 Max is billed at list times 1.25, so test one take at 480p first, and compare it with a Fabric render of the same line.
  • Keep the same voice and settings for both parts.
  • Never resubmit a paid job that looks slow; poll its status.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume