One portrait, three languages: three 6-second lip-sync jobs for $1.80

A greeting in English, Spanish and Polish from one still: three 6-second H3 Max 768p jobs cost $1.80 plus TTS. Steps, the language check and limits.

4 min readSume
All posts

Three H3 Max lip-sync jobs of 6 seconds each at 768p cost $1.80 in total, because each bills ceil(6) x $0.10 = $0.60. That is the lip-sync part of sending one portrait speaking English, Spanish and Polish; text-to-speech is billed separately, by characters. Each language needs its own audio file, so each is its own job, and every audio file has to land between 5 and 14.8 seconds.

The numbers

The rate comes from the repo's H3 Max lip-sync doc: the fal list of $0.08 per second at 768p times a 1.25 margin. The route reserves at admission, captures on completion and refunds on failure.

Three greetings from one still (read 2026-10-08)
LanguageAudio lengthBilled seconds768p cost
English6.0 s6$0.60
Spanish6.0 s6$0.60
Polish6.0 s6$0.60
Total18$1.80

Steps

Do the audio first and the video second. The same still goes into every job.

  • Write the greeting once and translate it. Spanish and Polish often run longer than English, so aim for lines that read in about 6 seconds in every language.
  • Create each audio file with TTS, setting language to match the voice. A voice whose language differs from the request returns 409 tts_voice_language_mismatch before any charge; confirm and retry with the same idempotency key only if you meant it.
  • Check the length of each file. If one is under 5 seconds, join two lines with the timeline audio endpoint; if it is over 14.8, shorten the text.
  • Submit each file to POST /v1/minimax/h3-max/lip-sync with the same image_url or avatar_handle, the audio URL and duration_seconds equal to the audio length.
  • Poll each job and download the three results.

Keep the three consistent

Use the same model, resolution and still for all three, so the faces match. The Sume packet guidance for lip sync is to use one model per run. If a line cannot reach 5 seconds, the docs suggest sending it to Fabric instead, but then that language would come from a different model, so lengthen the line first if you can.

Budgeting a campaign

If the greeting goes to ten markets instead of three, the lip-sync part is linear: ten 6-second jobs at 768p are 10 x $0.60 = $6.00. Resolution is the lever. The same ten at 480p cost 10 x $0.375 = $3.75, and at 1080p they cost $12.00. Because every job reserves at admission and refunds on failure, a batch that partly fails does not leave you paying for the failed jobs. Use a separate Idempotency-Key for each language, so that a retry of one language never creates a second charge for another.

What Sume does not do

Sume does not translate the audio of an existing clip or move the lips of a video you already shot. It starts from a still. It also does not choose a voice for the language for you: pick a voice whose language matches, and the check will tell you if it does not.

Speech length differs between languages, so there is no way to say in advance that a given sentence will land at exactly 6 seconds. Measure the audio, then set duration_seconds.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume