One-minute talking photo: five H3 Max lip sync clips, $6 at 768p
H3 Max lip sync takes 5-14.8 s of audio per clip. A 60-second spoken script is five 12 s clips: $3.75 at 480p, $6.00 at 768p, $12.00 at 1080p on Sume.

A one-minute spoken script becomes five H3 Max lip sync clips of about 12 seconds each, because one clip takes 5 to 14.8 seconds of audio. Billed at ceil(duration_seconds) per clip, 60 seconds of audio costs $3.75 at 480p, $6.00 at 768p and $12.00 at 1080p on Sume.
Sume's lip-sync route is POST /v1/minimax/h3-max/lip-sync. It takes a still (or a ready avatar) and a Sume-hosted audio file, and its price is the fal lip-sync list per second ($0.05, $0.08, $0.16) times the 1.25 house multiple.
Why five clips and not four
Four clips at the 14.8-second ceiling hold 59.2 seconds, which is short of a 60-second script. Five clips leave room to split on sentence boundaries, with each part between 5 and 14.8 seconds. The API refuses a duration_seconds outside that window with invalid_request; it never clamps. The docs explain why: the provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds, so a clamped reservation would produce a clip shorter than the audio and some speech would be lost.
That is the reason to split the script before you make the audio. Generate speech for each sentence group separately, so every file is a clean take that fits the window. The packet guidance in the docs says not to cut or re-synthesize the audio again to make it fit, because the waveform of a take should stay the same.
Cost by resolution
Each clip is billed for its whole seconds. Five clips of exactly 12 seconds is 60 billed seconds; if a clip is 12.3 seconds it bills 13. Plan your splits so that no clip lands just above a whole second.
| Resolution | List per second | Sume per second | Per 12 s clip | Five clips |
|---|---|---|---|---|
| 480p | $0.05 | $0.0625 | $0.75 | $3.75 |
| 768p (default) | $0.08 | $0.10 | $1.20 | $6.00 |
| 1080p | $0.16 | $0.20 | $2.40 | $12.00 |
Constraints to plan around
The audio must be hosted on Sume's media host and be 10 MB or less; the API rejects other sources with unsupported_audio_source or audio_too_large. The still needs a public HTTPS URL and an aspect ratio between 0.4 and 2.5, or you can use a ready avatar by id or handle. There is no 2K, and speed_tier is accepted and ignored.
The docs also advise using one lip-sync model for each run, because the model sets the assemble output.fps lock. Fabric is 25 fps, and the docs say to measure H3 Max on the first real clip. A minute made of five H3 Max clips should not mix in a Fabric clip.
- Split the script at sentence ends, targeting 10 to 13 seconds a part.
- Send one job per part with its own
Idempotency-Key. - Join the five outputs on a timeline in order; keep the same still for each.
Compare with the default talk model
Fabric is the default talk model on Sume, and the docs give its reservation as $0.94 for a 5-second clip at 720p, while a 5-second H3 Max lip sync clip at 768p reserves $0.50. Fabric covers audio outside the 5 to 14.8 second window, so a closing line of 3 seconds goes there, but then you mix two models in one run, and the fps lock above applies. If the script can be re-cut so that every part is 5 seconds or longer, keep it on one model.
Submitting the five jobs
Each clip is its own job of type avatar_image_to_video with the model minimax/h3-max/lip-sync. Send an Idempotency-Key for every part, and derive it from the script part so a retry cannot create a second paid job. The submit response carries usage.billable_amount_usd_micros; add the five quotes and compare the sum to the table before the first render finishes.
The reservation is taken at admit, captured on completion and refunded on failure, so a failed part does not leave you paying for it. Re-submit only the failed part, with the same still and audio, and keep the other four. If one part comes out with a weak mouth shape, a re-run is another take at the same price, since there is no seed control mentioned in the lip-sync docs.
Sources
Related posts
More in Use cases
- One motion clip, six colorways: Kling 3.0 Motion Control, $1.89 each
Kling 3.0 Motion Control bills $0.1575 per second. One 12-second motion clip applied to six product stills costs $11.34. The inputs and limits are inside.
- One narrator in many languages: a Sume TTS batch vs MAI-Voice-2.1
MAI-Voice-2.1 claims one voice across 23 languages. On Sume, you send the same voice.id with a language per request and audition each language.
- One Short in Korean and English: caption styles and rules
Posting a Korean and an English cut of your own AI Short? YouTube is silent on translations; Sume caption styles are not. Which style fits which script.
- One-take walkthrough prompts: store, kitchen, studio on Seedance 2.5
Three one-take walkthrough prompts for Seedance 2.5 (store aisle, restaurant kitchen, studio), each with a route, a pace and a 30 s price at 480p to 1080p.
Written by Sume