Sume TTS voice then H3 Max lip sync: 40 seconds in 3 slices
H3 Max lip-sync accepts 5 to 14.8 s of audio. Split a 40 second TTS script into three sentence-based slices at $4.00 on 768p, and what to check.
Split a 40 second line into three audio slices of 13, 14 and 13 seconds, lip-sync each slice with the MiniMax H3 Max route at 768p, and join the clips. At $0.10 per audio second, the 40 seconds of slices cost $4.00 (40 x 0.10). The route accepts 5 to 14.8 seconds of audio per call, so a 40 second take cannot go in as one request.
| Step | Detail | Cost |
|---|---|---|
| TTS with the avatar's voice | avatar_handle selector; wav output | Per the TTS rate card |
| Slice into 3 | Sentence segmentation, each 5 to 14.8 s | None |
| Lip sync at 768p | 40 s x 0.10 | $4.00 |
| Join clips | Timeline tools | Per the timeline rate card |
Getting sentence slices from TTS
The TTS description says that setting segmentation.mode=sentence returns gapless per-sentence audio slices, and that timestamps.words returns word timings. Sentence slices give you natural cut points, so a clip does not end mid-word. Ask for wav with pcm_s16le at 44100, which the same description names for use in avatar mux, so the file is ready for the lip-sync step.
Synthesized audio over 1200 seconds fails with tts_duration_exceeded and no credit capture, which is far above this plan.
Group sentences into windows
Add sentence lengths until the next one would pass 14.8 seconds, then start a new slice. A 40 second line might split as 13, 14 and 13 seconds. Keep every slice at 5 seconds or more, because shorter audio is outside the route's window; merge a short last sentence into the one before it.
- Slice A: 13.0 s x 0.10 = $1.30.
- Slice B: 14.0 s x 0.10 = $1.40.
- Slice C: 13.0 s x 0.10 = $1.30.
- Total 40.0 s = $4.00.
Checks before you ship
The route is billed per audio second and the output length follows the audio, so you do not control the cut. Compare the three clips for the same face and framing; each call renders separately and nothing in the docs promises identical lighting across calls. Use the same avatar handle for all three, and burn captions after the join with Video captions, not before.
When to skip this recipe
If the script is in English and under 60 seconds, the talking-video route does the planning, voice and rendering in one job, and it is simpler: 40 seconds on standard is 40 x 0.184 = $7.36. The TTS and lip-sync path costs less per second on its lip-sync step but needs you to manage the slices, the join and the captions. Choose it when you need a non-English voice, a voice you already approved, or exact control of timing.
Name the slice files with the order and the avatar handle, and keep the sentence text beside each file. When one slice needs a retake, you can rerun only that slice: 13 seconds at 768p is $1.30, instead of redoing all 40 seconds for $4.00.
Sources
Related posts
More in Sume Avatar 1.0
- Can a Sume avatar speak my own recording? Voice types explained
Avatar 1.0 talking video takes a script or per-scene voice blocks of type text or silence. What the docs list, what they do not, and how to plan a pause.
- Will an AI video model lip-sync a voice-over added afterwards?
No: Sume's docs say video models do not lip-sync to generated speech or a later voice-over. A talking face needs a still plus audio, or an avatar video.
- Face-swap clip on YouTube: label needed? Sume's beta caps at 15 s
YouTube asks for a label when a real person appears to say or do something they did not. Sume's face swap beta takes about 4 to 15 seconds; $3.68 max at plus.
- Face swap or avatar video? Pick by whether you already have footage
Sume's face swap Beta applies an avatar face to your own 4-15 second video; Avatar Video renders a new clip from a script, up to 60 seconds. How to choose.
Written by Sume