Sume TTS voice then H3 Max lip sync: 40 seconds in 3 slices

H3 Max lip-sync accepts 5 to 14.8 s of audio. Split a 40 second TTS script into three sentence-based slices at $4.00 on 768p, and what to check.

5 min readSume
All posts

Split a 40 second line into three audio slices of 13, 14 and 13 seconds, lip-sync each slice with the MiniMax H3 Max route at 768p, and join the clips. At $0.10 per audio second, the 40 seconds of slices cost $4.00 (40 x 0.10). The route accepts 5 to 14.8 seconds of audio per call, so a 40 second take cannot go in as one request.

Plan for a 40 second line, rates as of 2026-10-08
StepDetailCost
TTS with the avatar's voiceavatar_handle selector; wav outputPer the TTS rate card
Slice into 3Sentence segmentation, each 5 to 14.8 sNone
Lip sync at 768p40 s x 0.10$4.00
Join clipsTimeline toolsPer the timeline rate card

Getting sentence slices from TTS

The TTS description says that setting segmentation.mode=sentence returns gapless per-sentence audio slices, and that timestamps.words returns word timings. Sentence slices give you natural cut points, so a clip does not end mid-word. Ask for wav with pcm_s16le at 44100, which the same description names for use in avatar mux, so the file is ready for the lip-sync step.

Synthesized audio over 1200 seconds fails with tts_duration_exceeded and no credit capture, which is far above this plan.

Group sentences into windows

Add sentence lengths until the next one would pass 14.8 seconds, then start a new slice. A 40 second line might split as 13, 14 and 13 seconds. Keep every slice at 5 seconds or more, because shorter audio is outside the route's window; merge a short last sentence into the one before it.

  • Slice A: 13.0 s x 0.10 = $1.30.
  • Slice B: 14.0 s x 0.10 = $1.40.
  • Slice C: 13.0 s x 0.10 = $1.30.
  • Total 40.0 s = $4.00.

Checks before you ship

The route is billed per audio second and the output length follows the audio, so you do not control the cut. Compare the three clips for the same face and framing; each call renders separately and nothing in the docs promises identical lighting across calls. Use the same avatar handle for all three, and burn captions after the join with Video captions, not before.

When to skip this recipe

If the script is in English and under 60 seconds, the talking-video route does the planning, voice and rendering in one job, and it is simpler: 40 seconds on standard is 40 x 0.184 = $7.36. The TTS and lip-sync path costs less per second on its lip-sync step but needs you to manage the slices, the join and the captions. Choose it when you need a non-English voice, a voice you already approved, or exact control of timing.

Name the slice files with the order and the avatar handle, and keep the sentence text beside each file. When one slice needs a retake, you can rerun only that slice: 13 seconds at 768p is $1.30, instead of redoing all 40 seconds for $4.00.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume