One-minute talking photo: five H3 Max lip sync clips, $6 at 768p

H3 Max lip sync takes 5-14.8 s of audio per clip. A 60-second spoken script is five 12 s clips: $3.75 at 480p, $6.00 at 768p, $12.00 at 1080p on Sume.

5 min readSume
All posts

A one-minute spoken script becomes five H3 Max lip sync clips of about 12 seconds each, because one clip takes 5 to 14.8 seconds of audio. Billed at ceil(duration_seconds) per clip, 60 seconds of audio costs $3.75 at 480p, $6.00 at 768p and $12.00 at 1080p on Sume.

Sume's lip-sync route is POST /v1/minimax/h3-max/lip-sync. It takes a still (or a ready avatar) and a Sume-hosted audio file, and its price is the fal lip-sync list per second ($0.05, $0.08, $0.16) times the 1.25 house multiple.

Why five clips and not four

Four clips at the 14.8-second ceiling hold 59.2 seconds, which is short of a 60-second script. Five clips leave room to split on sentence boundaries, with each part between 5 and 14.8 seconds. The API refuses a duration_seconds outside that window with invalid_request; it never clamps. The docs explain why: the provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds, so a clamped reservation would produce a clip shorter than the audio and some speech would be lost.

That is the reason to split the script before you make the audio. Generate speech for each sentence group separately, so every file is a clean take that fits the window. The packet guidance in the docs says not to cut or re-synthesize the audio again to make it fit, because the waveform of a take should stay the same.

Cost by resolution

Each clip is billed for its whole seconds. Five clips of exactly 12 seconds is 60 billed seconds; if a clip is 12.3 seconds it bills 13. Plan your splits so that no clip lands just above a whole second.

60 seconds of audio as five 12 s lip sync clips, Sume billable (rate card 2026-09-22 x 1.25, docs read 2026-10-05)
ResolutionList per secondSume per secondPer 12 s clipFive clips
480p$0.05$0.0625$0.75$3.75
768p (default)$0.08$0.10$1.20$6.00
1080p$0.16$0.20$2.40$12.00

Constraints to plan around

The audio must be hosted on Sume's media host and be 10 MB or less; the API rejects other sources with unsupported_audio_source or audio_too_large. The still needs a public HTTPS URL and an aspect ratio between 0.4 and 2.5, or you can use a ready avatar by id or handle. There is no 2K, and speed_tier is accepted and ignored.

The docs also advise using one lip-sync model for each run, because the model sets the assemble output.fps lock. Fabric is 25 fps, and the docs say to measure H3 Max on the first real clip. A minute made of five H3 Max clips should not mix in a Fabric clip.

  • Split the script at sentence ends, targeting 10 to 13 seconds a part.
  • Send one job per part with its own Idempotency-Key.
  • Join the five outputs on a timeline in order; keep the same still for each.

Compare with the default talk model

Fabric is the default talk model on Sume, and the docs give its reservation as $0.94 for a 5-second clip at 720p, while a 5-second H3 Max lip sync clip at 768p reserves $0.50. Fabric covers audio outside the 5 to 14.8 second window, so a closing line of 3 seconds goes there, but then you mix two models in one run, and the fps lock above applies. If the script can be re-cut so that every part is 5 seconds or longer, keep it on one model.

Submitting the five jobs

Each clip is its own job of type avatar_image_to_video with the model minimax/h3-max/lip-sync. Send an Idempotency-Key for every part, and derive it from the script part so a retry cannot create a second paid job. The submit response carries usage.billable_amount_usd_micros; add the five quotes and compare the sum to the table before the first render finishes.

The reservation is taken at admit, captured on completion and refunded on failure, so a failed part does not leave you paying for it. Re-submit only the failed part, with the same still and audio, and keep the other four. If one part comes out with a weak mouth shape, a re-run is another take at the same price, since there is no seed control mentioned in the lip-sync docs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume