Lip-sync a 20-second line: H3 Max takes 14.8 s of audio, so split it
MiniMax H3 Max lip sync on Sume accepts 5 to 14.8 seconds of audio. How to cut a 20-second line into two takes, or use Fabric for a longer face-to-camera line.

The limit and the workaround
Sume's POST /v1/minimax/h3-max/lip-sync takes audio of 5 to 14.8 seconds, so a 20-second line must be cut into two takes of at least 5 seconds each, and each take run as its own job. Cut at a sentence boundary so nobody sees a word split in half, then join the two videos.
If the line is long and the speaker is a still photo, the other route is POST /v1/veed/fabric-1.0, which takes a still and audio of up to 10 MB with no 14.8-second ceiling.
Pick the route
The right route depends on what you start with: a video of a person, or a still.
| Route | Starts from | Audio rule | Notes |
|---|---|---|---|
| H3 Max lip sync | A video | 5 to 14.8 seconds | 480p, 768p or 1080p |
| VEED Fabric 1.0 | An accepted still | Sume media host audio, up to 10 MB | Still plus audio |
| Avatar talking video | A Sume avatar | 4 to 60 seconds | Avatar plus script |
Cutting a 20-second line in two
Write the script in two parts before you generate the voice, so each part is a whole thought of roughly 7 to 12 seconds. Generate each part as its own tts-1.0 job. This is better than cutting finished audio, because the voice finishes a sentence naturally. If you already have one 20-second file, use POST /v1/timeline-1.0/audio with operation: "split" to slice it into ranges at $0.01 flat, and choose a silent point between sentences.
Send each audio part with the same video to its own lip-sync job. Make sure each part is within 5 to 14.8 seconds, because a short tail under 5 seconds is refused. Then join the two clips in order with Sume's timeline. Expect a slight change of expression at the join; hide it with a cut-away or a caption. If the two takes must look continuous, start the second take from a frame close to where the first ended, and keep lighting, framing and resolution identical so the seam is hard to spot. Cost scales with the number of takes, so a 40-second line would need three or four, and at that length Fabric is simpler.
A rule from the docs
Sume's docs say that a face that talks is Fabric with an accepted still plus TTS, and that video models do not lip-sync to TTS. So do not expect a text-to-video model to sync a voice file you provide; use a lip-sync route for that.
- Check the
duration_secondsof each audio part before you submit. - Pay attention to price: H3 Max is billed at list times 1.25, so test one take at 480p first, and compare it with a Fabric render of the same line.
- Keep the same voice and settings for both parts.
- Never resubmit a paid job that looks slow; poll its status.
Sources
Related posts
More in Media tools
- LTX-2.5 48 fps option: conform frame rate on Sume
LTX-2.5 offers 24/25 fps or 48/50 fps. Sume's video catalog has no fps field, but Timeline output.fps conforms a clip to 24, 25, 30 or 60.
- Lyria 3.5 output: MP3 or WAV at 44.1 kHz, and what Sume returns
Google lists MP3 by default or WAV, 44.1 kHz stereo, with a SynthID watermark. Sume's music docs say the artifact is typically audio/mpeg. Check it in code.
- Lyria 3.5 prompts: [Verse] tags or [0:00-0:30] time ranges?
Google's Lyria 3.5 docs show [Verse], [Chorus] and [Bridge] tags. Sume's music docs show time-range markers. What each says, and how to test both on one brief.
- MAI-Voice-2.1 Flash narration for a video cut
Narration from any voice model can be the audio spine of a Sume Timeline cut. Import the file first, then render at $0.10 per output minute.
Written by Sume