Lip-sync a 2-minute monologue: split one TTS wav into H3 Max clips
MiniMax H3 Max lip-sync takes audio of 5 to 14.8 seconds. For a longer speech, render one TTS wav and cut it at sentence ends; the cost is worked below.

To lip-sync a speech longer than 14.8 seconds with H3 Max, render the whole script as one TTS wav with sentence segmentation, cut it into pieces of 5 to 14.8 seconds at sentence boundaries with timeline audio, and send each piece with the same still. A two-minute script becomes about nine clips.
This follows Sume's published limits: H3 Max lip-sync takes audio of 5 to 14.8 seconds per the Models overview, the TTS job returns timings in the API reference, and timeline audio splits a track by range.
Step 1: one TTS job with timings
Call POST /v1/tts-1.0/generate with your transcript, a voice, output_format set to a wav container, timestamps: {words: true} and segmentation: {mode: "sentence"}. The sentence mode needs word timestamps. The result carries gapless segments[] with start and end seconds. With wav output each segment may also carry its own audio_url when emit_audio is set, but slicing the single file keeps the voice flow continuous, so this plan uses the timings only.
- A transcript can be up to 20,000 characters, and audio past 1,200 seconds fails.
- Set
languagefor non-English text. - mp3 gives timings but no per-segment audio files.
Step 2: choose the cuts
Walk the segments and add sentences to a group until the next one would push the group past 14.8 seconds. Each group must be at least 5 seconds. A very short sentence can join its neighbour; one sentence longer than 14.8 seconds needs rewriting before you start, because you cannot cut inside it without a visible jump.
| Group | Length (s) | Billed seconds at H3 Max |
|---|---|---|
| 1 to 8 | 13.4 each (8 groups) | 14 each |
| 9 | 12.8 | 13 |
| Total | 120.0 | 125 |
Step 3: split and send
Use timeline_audio with operation: "split" and the ranges from step 2 (1 to 20 ranges per job, $0.01 flat, wav by default). Then call POST /v1/minimax/h3-max/lip-sync once per range with the same portrait and that audio. Check the route's audio field for whether it needs a Sume-hosted URL; the split outputs are hosted on Sume.
| Line | Math | Cost |
|---|---|---|
| TTS, 1,800 chars | 1.8 x $0.0475 = $0.0855, rounded to a cent | $0.09 |
| Split job | flat | $0.01 |
| H3 Max clips | 125 s x $0.10 | $12.50 |
| Total | $12.60 |
What to check before you pay for all nine
Run clip 1 first and look at it. If the face drifts or the audio edge clicks, fix it before the other eight. H3 Max bills per ceiling audio second, so a 13.4 second piece costs 14 seconds, and cutting pieces to land just under a whole second saves a little. Join the finished clips on a timeline in order.
When this plan is the wrong choice
If the monologue is under 14.8 seconds and over 5, skip the split and send the one TTS wav straight to the lip-sync route. If it is much longer than two minutes, consider cutting to B-roll between clips: nine consecutive talking clips from one still can look repetitive, and B-roll hides the seams. Fabric, the other talking-still route, takes audio up to 300 seconds per clip, so for long reads the question is price and look, not length; compare the 720p rate of $0.1875 per second with H3 Max at $0.10 for 768p before you choose.
Also keep rights in view: only animate a portrait you have the right to use, and only voice a person with their agreement.
| Route | Audio length | Rate |
|---|---|---|
| MiniMax H3 Max | 5 to 14.8 s | 768p $0.10 per ceiling second |
| VEED Fabric 1.0 | 1 to 300 s | 720p $0.1875 per second |
Sources
Related posts
More in Developers
- Is there a lipsync-1.0 endpoint on Sume? Old paths 404; use Fabric
Sume's old /v1/lipsync-1.0 paths return 404 and its model ids return model_not_found. Send the same still and audio to veed/fabric-1.0 or H3 Max lip-sync.
- List Sume's music and TTS router engines in Python before deploy
A short Python script reads the music and TTS router model lists so a deploy check can confirm your pinned engine ids and fixed prices still exist.
- Does Sume accept lowercase 4k as a video resolution? Yes
Since fab7d59af, lowercase 4k is an alias on every Sume video row that lists 4K, on both /v1/video-router/generate and /v1/videos.
- Lyria 3.5 is single-turn: what that means for Sume Music retries
Google says Lyria music is single-turn, with no iterative editing and varying results per call. On Sume that means a new take per prompt: no seed, no duration.
Written by Sume