Fit TTS audio to the 5 to 14.8 second H3 Max lip-sync window

Sume's H3 Max lip-sync accepts audio of 5 to 14.8 seconds and never clamps. Plan scripts of about 63 to 197 characters, and send anything else to Fabric.

4 min readSume
All posts

Sume's MiniMax H3 Max lip-sync route takes a still and an audio file and only accepts audio between 5 and 14.8 seconds. If your duration_seconds is outside that range the API returns invalid_request; it never clamps. So the fit has to happen before you call it: write the script to land in the window, generate the speech, and send anything that falls outside to Fabric, which stays the default talking model.

What the route enforces

From Sume's H3 Max lip-sync reference: POST /v1/minimax/h3-max/lip-sync, with the public model id minimax/h3-max/lip-sync.

H3 Max lip-sync input rules, Sume reference
FieldRule
audio_urlSume-hosted, at most 10 MB
duration_secondsRequired, 5 to 14.8; used as the reservation basis
Imageimage_url (public HTTPS, aspect ratio 0.4 to 2.5) or a ready avatar by avatar_id or avatar_handle, exactly one
resolution480p, 768p (default) or 1080p; no 2K
Out of rangeRejected with invalid_request; never clamped

Why the API does not clamp

The reference explains the failure it avoids: the provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds. If Sume clamped the reservation, the rendered clip would be shorter than your audio and some speech would be lost without any error. It prefers a clear rejection. The reference also advises, though the API does not enforce it, never to cut or re-synthesize the audio to make it fit. Segments outside the window go to Fabric instead.

Planning a script that lands in the window

The TTS Router does not take a duration, so length is set by your text. Cartesia's pricing page says a minute of audio takes 750 to 800 credits, with one credit per character. That is 12.5 to 13.3 characters per second, which gives a rough planning range. It is a billing ratio, not a promise: speaking pace, language and punctuation change it, so always read the real duration of the generated file.

Character budget for the window, arithmetic on Cartesia's 750 to 800 characters per minute
Target audioCharacters at 12.5 per secondCharacters at 13.3 per second
5 s (lower limit)about 63about 67
10 sabout 125about 133
14.8 s (upper limit)about 185about 197

A workflow that respects the window

Talking shots use TTS first and then a lip-sync model, never a video model. Here is the order that avoids a rejection.

  • Split the script into lines of one to two sentences, aiming for about 80 to 160 characters.
  • Generate speech with POST /v1/tts-router/generate or TTS 1.0 and keep the Sume-hosted audio URL.
  • Measure the duration of each audio file. Do not rely on the character estimate.
  • Between 5 and 14.8 seconds: send to H3 Max with that duration, rounded to hundredths.
  • Outside the window: use Fabric for that line. Do not trim the audio.
  • Use one lip-sync model per run, because the model sets the output frame rate (Fabric is 25 fps; H3 Max must be measured on a first clip).

What it costs

Billing uses ceil(duration_seconds) times the per-second rate, so a 5.84 second clip bills as 6 seconds. At 768p that is 6 x $0.08 x 1.25 = $0.60, before the platform fee. Padding a line to 14.8 seconds costs $1.50 at 768p (15 x $0.08 x 1.25), so shorter lines are cheaper. For clips with a stored avatar, see the Avatar 1.0 docs.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume