Write a lip-sync script in 5 to 14.8 second beats (H3 Max voice-over)

MiniMax H3 Max lip-sync on Sume takes audio of 5 to 14.8 seconds. Write the script in beats that fit the window, then voice each beat; the price per beat.

5 min readSume
All posts

To make a talking clip with MiniMax H3 Max lip-sync on Sume, write the script in beats that each land between 5 and 14.8 seconds when spoken, voice each beat as its own audio file, and submit one job per beat. The window is a hard limit of the route: the API refuses a duration_seconds outside 5 to 14.8, and it never clamps. Fixing the length at the script stage is cheaper than discovering it at the submit stage.

The reason for the strictness is in the Sume docs. The provider rejects audio under 5 seconds and silently clips anything after 14.8 seconds, so a clamped reservation would render a clip shorter than your audio and lose some of the speech. Sume refuses the request instead of letting that happen.

A practical habit is to keep a small sheet next to the script with one row per beat: the text, the measured length of the voiced file, the billed seconds and the resolution. When a beat is rewritten, the row is updated and the bill is known before any job is submitted. That sheet is also the shot list you assemble the final cut from, so nothing is lost by writing it.

The limits that shape a beat

Everything below comes from the lip-sync route documentation. The route takes a still image (or a ready avatar), a Sume-hosted audio file and the audio length, and returns a talking clip whose length is set by the audio.

MiniMax H3 Max lip-sync request limits (Sume docs, read 2026-10-07)
ItemLimit
RoutePOST /v1/minimax/h3-max/lip-sync
Audio length5 to 14.8 seconds
duration_secondsRequired, 5 to 14.8, refused outside the range
Audio sourceaudio_url on the Sume media host, up to 10 MB
Resolution480p, 768p (default) or 1080p; no 2K
StillPublic HTTPS, aspect ratio 0.4 to 2.5, or avatar_id

Writing the script in beats

Draft the script as one idea per beat. A beat is a sentence or two that you could say in a single breath of the on-screen speaker, and it ends where a cut would be natural. Read each beat aloud with a timer, or voice it and measure the file, and keep the beats that land between 5 and 14.8 seconds. A beat under 5 seconds either gets merged with its neighbor or goes to Fabric, which stays the default talk model for segments outside the H3 Max window.

A longer speech is several beats, cut on the timeline. Do not stretch or trim the audio to make it fit. The packet guidance in the docs says never to cut or synthesize the audio again to fit the window, because that changes the waveform. Rewrite the beat instead, and voice it again.

  • Under 5 s: merge it with the next beat, or send it to Fabric.
  • 5 to 14.8 s: one H3 Max lip-sync job.
  • Over 14.8 s: split the text at a sentence break and voice two beats.
  • Use one lip-sync model for each run, because the model sets the frame rate you assemble to.

Submit one beat

The request body matches the Fabric body. This example uses a still image and a Sume-hosted audio file for a 7.3 second beat at the default resolution. The duration_seconds value is the length of your audio file, not a guess.

curl -X POST "https://api.sume.com/v1/minimax/h3-max/lip-sync" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Idempotency-Key: beat-03-v1" \
  -H "Content-Type: application/json" \
  -d '{
    "image_url": "https://example.com/presenter.png",
    "audio_url": "https://media.sume.com/your-audio-file",
    "duration_seconds": 7.3,
    "resolution": "768p",
    "mode": "async"
  }'

What each beat costs

The price is the whole seconds of duration_seconds rounded up, times the per-second rate, times the 1.25 house margin. At 768p the rate is $0.10 per second, which is a $0.08 provider list plus the margin. A 7.3 second beat bills 8 seconds, so $0.80. A beat of 5.1 seconds bills 6 seconds, which is why the stored note on ceil rounding is worth reading before you tune beat lengths.

Plan the beats so they end just under a whole second where you can: a beat of 9.9 seconds bills 10 seconds, and one of 10.1 bills 11. The check on the voiced file is also covered in checking TTS length before H3 Max lip-sync.

Bill for one beat by length and resolution (provider list x 1.25 per second, whole seconds; Sume docs, read 2026-10-07)
Beat lengthBilled seconds480p ($0.0625/s)768p ($0.10/s)1080p ($0.20/s)
5.0 s5$0.3125$0.50$1.00
7.3 s8$0.50$0.80$1.60
10.0 s10$0.625$1.00$2.00
14.8 s15$0.9375$1.50$3.00

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume