Lip-sync a 2-minute monologue: split one TTS wav into H3 Max clips

MiniMax H3 Max lip-sync takes audio of 5 to 14.8 seconds. For a longer speech, render one TTS wav and cut it at sentence ends; the cost is worked below.

6 min readSume
All posts

To lip-sync a speech longer than 14.8 seconds with H3 Max, render the whole script as one TTS wav with sentence segmentation, cut it into pieces of 5 to 14.8 seconds at sentence boundaries with timeline audio, and send each piece with the same still. A two-minute script becomes about nine clips.

This follows Sume's published limits: H3 Max lip-sync takes audio of 5 to 14.8 seconds per the Models overview, the TTS job returns timings in the API reference, and timeline audio splits a track by range.

Step 1: one TTS job with timings

Call POST /v1/tts-1.0/generate with your transcript, a voice, output_format set to a wav container, timestamps: {words: true} and segmentation: {mode: "sentence"}. The sentence mode needs word timestamps. The result carries gapless segments[] with start and end seconds. With wav output each segment may also carry its own audio_url when emit_audio is set, but slicing the single file keeps the voice flow continuous, so this plan uses the timings only.

  • A transcript can be up to 20,000 characters, and audio past 1,200 seconds fails.
  • Set language for non-English text.
  • mp3 gives timings but no per-segment audio files.

Step 2: choose the cuts

Walk the segments and add sentences to a group until the next one would push the group past 14.8 seconds. Each group must be at least 5 seconds. A very short sentence can join its neighbour; one sentence longer than 14.8 seconds needs rewriting before you start, because you cannot cut inside it without a visible jump.

Example grouping of a 120 second read into nine clips (illustrative arithmetic)
GroupLength (s)Billed seconds at H3 Max
1 to 813.4 each (8 groups)14 each
912.813
Total120.0125

Step 3: split and send

Use timeline_audio with operation: "split" and the ranges from step 2 (1 to 20 ranges per job, $0.01 flat, wav by default). Then call POST /v1/minimax/h3-max/lip-sync once per range with the same portrait and that audio. Check the route's audio field for whether it needs a Sume-hosted URL; the split outputs are hosted on Sume.

Cost of the plan (Sume rates: TTS $0.0475 per 1,000 chars, split $0.01, H3 Max 768p $0.10 per audio second; checked 2026-10-10)
LineMathCost
TTS, 1,800 chars1.8 x $0.0475 = $0.0855, rounded to a cent$0.09
Split jobflat$0.01
H3 Max clips125 s x $0.10$12.50
Total$12.60

What to check before you pay for all nine

Run clip 1 first and look at it. If the face drifts or the audio edge clicks, fix it before the other eight. H3 Max bills per ceiling audio second, so a 13.4 second piece costs 14 seconds, and cutting pieces to land just under a whole second saves a little. Join the finished clips on a timeline in order.

When this plan is the wrong choice

If the monologue is under 14.8 seconds and over 5, skip the split and send the one TTS wav straight to the lip-sync route. If it is much longer than two minutes, consider cutting to B-roll between clips: nine consecutive talking clips from one still can look repetitive, and B-roll hides the seams. Fabric, the other talking-still route, takes audio up to 300 seconds per clip, so for long reads the question is price and look, not length; compare the 720p rate of $0.1875 per second with H3 Max at $0.10 for 768p before you choose.

Also keep rights in view: only animate a portrait you have the right to use, and only voice a person with their agreement.

Two lip-sync routes (Sume docs, checked 2026-10-10)
RouteAudio lengthRate
MiniMax H3 Max5 to 14.8 s768p $0.10 per ceiling second
VEED Fabric 1.01 to 300 s720p $0.1875 per second

Sources

Related posts

More in Developers

All Developers posts

Written by Sume