MiniMax H3 native audio vs TTS plus lip sync for a 10-second line
MiniMax H3 makes stereo audio with the clip. For an exact spoken line, Sume prices TTS plus Fabric or H3 Max lip sync: $0.75 to $2.01 for 10 seconds.

MiniMax H3 shipped on July 31 with 15-second clips, 2K hosting and native stereo audio, as the Magic Hour tracker reports (read 2026-10-07). On Sume, minimax-h3 and minimax-h3-max both generate audio with the clip. The question for a spoken line is whether to let the model make the voice or to record the words yourself with text to speech and then drive a face with them.
Sume's own rule, from the model docs, is plain: a shot where a person speaks on camera is Fabric with an accepted still plus TTS, because video models do not lip-sync to generated TTS or a later voice-over. So the choice is between an H3 clip whose audio you accept as it comes, and a TTS plus lip-sync pipeline whose words you control.
Price of a 10-second spoken line on Sume
All rates are the provider list times the 1.25 house margin, read from the repo's catalog on 2026-10-07. The TTS line assumes about 150 characters, which is $0.0071 and rounds up to $0.01.
| Route | Rate | 10 seconds | You control the words |
|---|---|---|---|
| MiniMax H3, 768p, native audio | $0.075 per second | $0.75 | No, the model writes the sound |
| MiniMax H3 Max, 768p, native stereo audio | $0.10 per second | $1.00 | No |
| TTS + Fabric 480p | $0.0475 per 1,000 characters + $0.10 per audio second | $1.01 | Yes |
| TTS + Fabric 720p | $0.0475 per 1,000 characters + $0.1875 per audio second | $1.89 | Yes |
| TTS + H3 Max lip sync, 768p | $0.0475 per 1,000 characters + $0.10 per audio second | $1.01 | Yes |
| TTS + H3 Max lip sync, 1080p | $0.0475 per 1,000 characters + $0.20 per audio second | $2.01 | Yes |
Limits that change the answer
- H3 and H3 Max accept 5 to 15 seconds; H3 Max reads 480p, 768p or 1080p.
- H3 Max lip sync takes audio of 5 to 14.8 seconds. Sume refuses both edges instead of clamping, so a 4-second line must be padded.
- Fabric takes up to 300 seconds of audio, Sume-hosted, under 10 MB, so it suits long talking blocks.
- A lip-sync clip's length follows the audio, so the price is audio seconds times the rate, rounded up per second.
When to pick which
Pick native audio when the sound is atmosphere: footsteps, a crowd, music under a product shot. The 768p H3 clip at $0.75 for ten seconds is the cheapest line in the table, and the sound is part of the price. Pick TTS plus a lip-sync model when a sentence has to be right, such as a price, a brand name or a legal line. The extra cost is about a cent of speech, and the lip-sync seconds cost what the video seconds would have.
Whichever route you use, listen to the result before it goes into a longer cut; the repo's tools describe the speech and sound as outputs to check, not guaranteed values.
Sources
Related posts
More in Models
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
- MiniMax H3 clip audio: keep it or detach it for a Short
MiniMax H3 is reported to make native stereo audio. Detach it for $0.01 to mix it separately. Prices for H3 768p and H3 Max.
- Minimum clip length by Sume video model: 2, 3, 4 or 5 seconds
Wan 3.0 starts at 2 s, Gemini Omni Flash 1.1 at 3 s, Seedance and Genjutsu at 4 s, MiniMax H3 and Recast at 5 s. Full min and max table for a Sora port.
- Music Router image_url: public HTTPS only, and null to clear it
Sume music image_url must be a public HTTPS image, and null clears it on a reused request object. When a mood still helps and when the prompt does more.
Written by Sume