AI lip sync API cost per second: H3 Max 480p, 768p, 1080p vs Fabric
On Sume, H3 Max lip-sync is $0.0625, $0.10 or $0.20 per audio second by resolution; VEED Fabric is $0.10 or $0.1875. A 14.8 s clip costs $0.94 to $3.00.

A lip-sync clip on Sume is billed per second of the audio you send, rounded up. MiniMax H3 Max lip-sync is $0.0625 per audio second at 480p, $0.10 at 768p (the default) and $0.20 at 1080p. VEED Fabric 1.0, the other audio-driven route, is $0.10 at 480p and $0.1875 at 720p. A full-length 14.8 second H3 Max clip bills as 15 seconds, so it costs $0.9375, $1.50 or $3.00 depending on resolution.
These are Sume's own published rates from its catalog and OpenAPI. This post covers the numbers, the audio-length window that limits H3 Max, and how to pick the route that costs least for the clip you have.
What are the two routes and their windows?
Both routes take a still image (a public image_url, or a ready avatar through avatar_id or avatar_handle) plus Sume-hosted audio and return a clip whose mouth follows the audio. Output length follows the audio.
The body is the same for both; the limits differ. H3 Max lip-sync, at POST /v1/minimax/h3-max/lip-sync, needs audio of 5 to 14.8 seconds, because the provider rejects shorter audio and silently clips longer audio, so Sume refuses requests outside the window. Fabric, at POST /v1/veed/fabric-1.0, accepts up to 300 seconds. Both require the audio to be on Sume's media host and under 10 MB, which is typically a text-to-speech segment.
What does one clip cost?
Cost is whole audio seconds times the rate for the resolution you pick. The H3 Max maximum of 14.8 seconds bills as 15.
| Audio length | H3 Max 480p ($0.0625) | H3 Max 768p ($0.10) | H3 Max 1080p ($0.20) | Fabric 720p ($0.1875) |
|---|---|---|---|---|
| 5 s | $0.3125 | $0.50 | $1.00 | $0.9375 |
| 10 s | $0.625 | $1.00 | $2.00 | $1.875 |
| 14.8 s (15 billed) | $0.9375 | $1.50 | $3.00 | $2.8125 |
| 60 s | not offered | not offered | not offered | $11.25 |
Which route is cheaper for my clip?
Inside the 5 to 14.8 second window, H3 Max at 768p is $0.10 per second against Fabric's $0.1875 at 720p, which is 47 percent less. Sume's own wording calls H3 Max the explicit alternative to Fabric inside that audio window. Resolution is the thing to compare, not just price: H3 Max reaches 1080p, Fabric tops out at 720p, and 1080p H3 Max at $0.20 costs slightly more per second than Fabric 720p.
Outside the window the choice is made for you. A 40 second spoken line has to go through Fabric, or be cut into pieces of 14.8 seconds or less and run through H3 Max, which gives you several clips to join and several reservations to pay. Split the line at sentence boundaries if you do that.
- Under 5 seconds of audio: H3 Max refuses it; use Fabric.
- 5 to 14.8 seconds: H3 Max at 480p or 768p is the cheapest published per-second rate.
- Over 14.8 seconds: Fabric, up to 300 seconds, at $0.10 (480p) or $0.1875 (720p).
- The
speed_tierfield is accepted by H3 Max and ignored; it does not change the price.
What does the speech itself cost?
The audio has to exist before the lip-sync job, and on Sume it usually comes from text to speech at $0.0475 per 1,000 characters. A 300-character line is $0.01425 of speech, so for short clips the video dominates the bill. A 10 second 768p clip is $1.00 of lip-sync plus about a cent and a half of voice. Ten such clips for a product explainer series come to roughly $10.14 before any joining step, and the same ten at 480p come to $6.39. Budget the video first and treat the voice as rounding.
How is the money held, and what fails early?
duration_seconds is required on both routes and is what the reservation uses: ceiling of the seconds times the resolution rate times 1.25 for H3 Max. Declare the real audio length. Sume reserves at submit, captures on completion, and releases or refunds on a failed job or a cancel before generation starts.
If the balance cannot cover the reservation, the response is 402 insufficient_credits and no job starts. A duration_seconds outside 5 to 14.8 on H3 Max is refused rather than clamped, so you are never billed for a window the provider would have cut. A full queue returns 429 queue_full and releases the reservation for that failed attempt. Use GET /v1/usage?job_id=... afterwards to read what that one clip cost.
Whichever model you pick, read the balance first and compare it with the rounded figure, since a short shortfall returns a 402 before any job starts.
For a budget, write down the audio length first. Because lip-sync audio is billed in whole seconds, a 9.3 second voiceover is billed as 10 seconds, and a clip near the 14.8 second limit is billed as 15. Compare that figure at your resolution against the Fabric figure for the same audio, then pick the cheaper model that meets your quality bar. Re-read the catalog before a large batch, since the live catalog is the source if a rate changes after the date above.
Sources
Related posts
More in Pricing
- AI virtual try-on cost per image: GPT Image 2.5 on Sume
What one try-on still costs on Sume: the documented token rates, why quality auto reserves the max price, and a script that reads the live endpoint pricing.
- AI voice agent cost per minute: speech-to-text, LLM and TTS stack
A dated cost sheet for the October 2026 voice stack: MAI-Transcribe-2-Streaming, Mercury Voice, MAI-Voice-2.1-Flash, and Sume's file STT and TTS, per minute.
- Change a person in a video with AI: Sume Recast price, not free
Changing a person in a video with Sume's H3 Max Recast is paid per source second: $0.375 at 768p, $0.5625 at 1080p. Costs for 5 to 30 second clips.
- Cost per row of a Sume bulk run: add up debited_usd_micros
Read usage.debited_usd_micros on each child receipt, wait for final to be true, and treat null as unknown. Why billable_amount alone understates a bulk run.
Written by Sume