Fit TTS audio to the 5 to 14.8 second H3 Max lip-sync window
Sume's H3 Max lip-sync accepts audio of 5 to 14.8 seconds and never clamps. Plan scripts of about 63 to 197 characters, and send anything else to Fabric.

Sume's MiniMax H3 Max lip-sync route takes a still and an audio file and only accepts audio between 5 and 14.8 seconds. If your duration_seconds is outside that range the API returns invalid_request; it never clamps. So the fit has to happen before you call it: write the script to land in the window, generate the speech, and send anything that falls outside to Fabric, which stays the default talking model.
What the route enforces
From Sume's H3 Max lip-sync reference: POST /v1/minimax/h3-max/lip-sync, with the public model id minimax/h3-max/lip-sync.
| Field | Rule |
|---|---|
audio_url | Sume-hosted, at most 10 MB |
duration_seconds | Required, 5 to 14.8; used as the reservation basis |
| Image | image_url (public HTTPS, aspect ratio 0.4 to 2.5) or a ready avatar by avatar_id or avatar_handle, exactly one |
resolution | 480p, 768p (default) or 1080p; no 2K |
| Out of range | Rejected with invalid_request; never clamped |
Why the API does not clamp
The reference explains the failure it avoids: the provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds. If Sume clamped the reservation, the rendered clip would be shorter than your audio and some speech would be lost without any error. It prefers a clear rejection. The reference also advises, though the API does not enforce it, never to cut or re-synthesize the audio to make it fit. Segments outside the window go to Fabric instead.
Planning a script that lands in the window
The TTS Router does not take a duration, so length is set by your text. Cartesia's pricing page says a minute of audio takes 750 to 800 credits, with one credit per character. That is 12.5 to 13.3 characters per second, which gives a rough planning range. It is a billing ratio, not a promise: speaking pace, language and punctuation change it, so always read the real duration of the generated file.
| Target audio | Characters at 12.5 per second | Characters at 13.3 per second |
|---|---|---|
| 5 s (lower limit) | about 63 | about 67 |
| 10 s | about 125 | about 133 |
| 14.8 s (upper limit) | about 185 | about 197 |
A workflow that respects the window
Talking shots use TTS first and then a lip-sync model, never a video model. Here is the order that avoids a rejection.
- Split the script into lines of one to two sentences, aiming for about 80 to 160 characters.
- Generate speech with
POST /v1/tts-router/generateor TTS 1.0 and keep the Sume-hosted audio URL. - Measure the duration of each audio file. Do not rely on the character estimate.
- Between 5 and 14.8 seconds: send to H3 Max with that duration, rounded to hundredths.
- Outside the window: use Fabric for that line. Do not trim the audio.
- Use one lip-sync model per run, because the model sets the output frame rate (Fabric is 25 fps; H3 Max must be measured on a first clip).
What it costs
Billing uses ceil(duration_seconds) times the per-second rate, so a 5.84 second clip bills as 6 seconds. At 768p that is 6 x $0.08 x 1.25 = $0.60, before the platform fee. Padding a line to 14.8 seconds costs $1.50 at 768p (15 x $0.08 x 1.25), so shorter lines are cheaper. For clips with a stored avatar, see the Avatar 1.0 docs.
Sources
Related posts
More in Media tools
- MiniMax H3 takes 9 image references: a slot plan for one product clip
Hailuo 3.0 (MiniMax H3) accepts up to 9 reference images per clip. A 9-slot plan for a product ad, and the request to send it on Sume.
- Minor technical edits vs real edits on YouTube Shorts
YouTube told creators to bring a voice, not minor technical or template changes. What that means for crops, filters and speed tweaks, and what real edits add.
- Mirror-floor reflection under an AI product cutout, built in Pillow
Flip a transparent product PNG from Sume, fade it with a gradient mask, and stack it under the original for a glossy floor reflection. Code and fade settings.
- Kling motion control reference over 30 seconds: trim it first
Sume's Kling motion control takes a duration_seconds of 1 to 30. For a longer dance or walk clip, cut a 30-second range with video trim, then submit that.
Written by Sume