One portrait, three languages: three 6-second lip-sync jobs for $1.80
A greeting in English, Spanish and Polish from one still: three 6-second H3 Max 768p jobs cost $1.80 plus TTS. Steps, the language check and limits.

Three H3 Max lip-sync jobs of 6 seconds each at 768p cost $1.80 in total, because each bills ceil(6) x $0.10 = $0.60. That is the lip-sync part of sending one portrait speaking English, Spanish and Polish; text-to-speech is billed separately, by characters. Each language needs its own audio file, so each is its own job, and every audio file has to land between 5 and 14.8 seconds.
The numbers
The rate comes from the repo's H3 Max lip-sync doc: the fal list of $0.08 per second at 768p times a 1.25 margin. The route reserves at admission, captures on completion and refunds on failure.
| Language | Audio length | Billed seconds | 768p cost |
|---|---|---|---|
| English | 6.0 s | 6 | $0.60 |
| Spanish | 6.0 s | 6 | $0.60 |
| Polish | 6.0 s | 6 | $0.60 |
| Total | 18 | $1.80 |
Steps
Do the audio first and the video second. The same still goes into every job.
- Write the greeting once and translate it. Spanish and Polish often run longer than English, so aim for lines that read in about 6 seconds in every language.
- Create each audio file with TTS, setting
languageto match the voice. A voice whose language differs from the request returns 409tts_voice_language_mismatchbefore any charge; confirm and retry with the same idempotency key only if you meant it. - Check the length of each file. If one is under 5 seconds, join two lines with the timeline audio endpoint; if it is over 14.8, shorten the text.
- Submit each file to
POST /v1/minimax/h3-max/lip-syncwith the sameimage_urloravatar_handle, the audio URL andduration_secondsequal to the audio length. - Poll each job and download the three results.
Keep the three consistent
Use the same model, resolution and still for all three, so the faces match. The Sume packet guidance for lip sync is to use one model per run. If a line cannot reach 5 seconds, the docs suggest sending it to Fabric instead, but then that language would come from a different model, so lengthen the line first if you can.
Budgeting a campaign
If the greeting goes to ten markets instead of three, the lip-sync part is linear: ten 6-second jobs at 768p are 10 x $0.60 = $6.00. Resolution is the lever. The same ten at 480p cost 10 x $0.375 = $3.75, and at 1080p they cost $12.00. Because every job reserves at admission and refunds on failure, a batch that partly fails does not leave you paying for the failed jobs. Use a separate Idempotency-Key for each language, so that a retry of one language never creates a second charge for another.
What Sume does not do
Sume does not translate the audio of an existing clip or move the lips of a video you already shot. It starts from a still. It also does not choose a voice for the language for you: pick a voice whose language matches, and the check will tell you if it does not.
Speech length differs between languages, so there is no way to say in advance that a given sentence will land at exactly 6 seconds. Measure the audio, then set duration_seconds.
Sources
Related posts
More in Media tools
- One TTS job per sentence, then a gapless concat: the Sume recipe
Generate narration one sentence at a time with tts_create, then join the takes with timeline audio. Fix one line without redoing the rest.
- Price plate over a product clip: Timeline compose overlay, $0.02
Sume Timeline compose overlay puts a price plate image on a product clip for $0.02. Width_ratio 0.05 to 1, margin_ratio 0 to 0.45, and the stack-key error.
- Do punch and tiktok-green captions accept design overrides? No
In Sume video captions, punch and tiktok-green do not support design: they render on a path that reads none of the tokens. Use another style to tune.
- Brand-color captions: punch and tiktok-green skip design
Sume's docs say punch and tiktok-green read none of the design tokens. To change a caption color per request, use slam or a Hangul style. Token groups.
Written by Sume