Lip-sync and motion control cost per clip: H3 Max, Fabric, Kling
Sume prices H3 Max lip-sync at $0.0625 to $0.20 per audio second, Fabric at $0.1875 and Kling 3.0 Motion Control at $0.1575. A 10-second clip is $0.63 to $2.00.

Three Sume tools turn existing media into new video. MiniMax H3 Max lip-sync costs $0.0625, $0.10 or $0.20 per audio second at 480p, 768p or 1080p. VEED Fabric 1.0 costs $0.1875 per audio second at 720p. Kling 3.0 Motion Control costs $0.1575 per output second. A 10-second clip is $0.625 to $2.00, and 100 clips are $62.50 to $200.00.
How each is billed
H3 Max lip-sync bills by ceil audio second: the Sume docs give a fal list of $0.05, $0.08 and $0.16 at 480p, 768p and 1080p, times the 1.25 house margin. The audio must be 5 to 14.8 seconds, Sume-hosted and under 10 MB. The docs' own example: a 5-second 768p clip reserves $0.50 (5 x $0.08 x 1.25).
Fabric is $0.1875 per audio second at 720p, up to 300 seconds. The docs' example: 5 seconds reserves $0.94 (5 x $0.1875 = $0.9375, rounded up to the cent). Kling 3.0 Motion Control is ceil(motion video second) x $0.126 list x 1.25 = $0.1575, for 1 to 30 seconds.
| Tool | Per second | 10 s clip | 100 clips |
|---|---|---|---|
| MiniMax H3 Max lip-sync, 480p | $0.0625 | $0.625 | $62.50 |
| MiniMax H3 Max lip-sync, 768p | $0.10 | $1.00 | $100.00 |
| MiniMax H3 Max lip-sync, 1080p | $0.20 | $2.00 | $200.00 |
| VEED Fabric 1.0, 720p | $0.1875 | $1.875 | $187.50 |
| Kling 3.0 Motion Control | $0.1575 | $1.575 | $157.50 |
A vendor reference point
fal's pricing page lists Kling Video v3 Image to Video [Pro] at $0.14 per second (read 2026-10-07). That is Kling's image-to-video product, not Motion Control, which the Sume docs list at $0.126 per second from fal in August. Different products, so do not subtract the two numbers. The page also lists the H3 Max text, image and reference-to-video models at $0.05 per second (read 2026-10-07); that is a different product from lip-sync, so I do not compare it with the lip-sync rate.
Choosing among them
Use lip-sync when the video exists and the voice is new: you pay by audio length, which is capped at 14.8 seconds. Use Fabric when you have a still and a voice track and want a talking clip, with up to 300 seconds. Use Motion Control to drive a character image with the motion of a reference video, with output length following the driving clip.
The 480p lip-sync draft at $62.50 per 100 ten-second clips against $200.00 at 1080p is a 3.2x spread. Draft cheaply, then re-render only the keepers.
Limits worth knowing
Fabric's audio must be hosted on Sume and under 10 MB, the same as lip-sync. Sume's separate Avatar Fabric route is experimental, priced at $0.3024 per second at 720p, and the docs say not to build production on it. These tools are metered per job.
A localization example
Take one 12-second English video and dub it into 5 languages with H3 Max lip-sync at 768p. Each dub is 12 x $0.10 = $1.20, so 5 languages cost $6.00. For 20 source videos that is $120.00. At 480p it falls to 12 x $0.0625 = $0.75 per dub, $75.00 for the 100 dubs, and at 1080p it rises to $2.40 per dub, $240.00.
The audio is separate. You need the translated voice track, which you can produce with TTS at $0.0475 per 1,000 characters; a 12-second line of about 160 characters is $0.0076. The voice is a rounding error next to the video, so the resolution choice is the real lever.
Sources
Related posts
More in Models
- Music Router prompt limit: 5,000 characters, and what to put in them
Sume Music Router prompts run 1 to 5,000 characters, with no duration field or negative prompt; length goes in the text. A brief budget and a checker.
- MAI-Transcribe-2-Streaming 0.13 s to final: when the clock starts
The 0.13 second figure is measured from end of speech found by a VAD, and partials arrive in about 100 ms. Why neither number is the wait for a file transcript.
- MAI-Transcribe-2-Streaming 2.5% WER: what the test audio mix is
The 2.5% word error rate comes from a chunked-streaming index: 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. Here is how that maps to a recorded call.
- MAI-Voice-2.1 has 97 voices: en-US has 7, es-ES and en-AU have 1 each
Microsoft's MAI-Voice-2.1 doc lists 97 prebuilt voices over 28 locales. We counted per locale: some have seven choices, others only one.
Written by Sume