MiniMax H3 native audio vs TTS plus lip sync for a 10-second line

MiniMax H3 makes stereo audio with the clip. For an exact spoken line, Sume prices TTS plus Fabric or H3 Max lip sync: $0.75 to $2.01 for 10 seconds.

5 min readSume
All posts

MiniMax H3 shipped on July 31 with 15-second clips, 2K hosting and native stereo audio, as the Magic Hour tracker reports (read 2026-10-07). On Sume, minimax-h3 and minimax-h3-max both generate audio with the clip. The question for a spoken line is whether to let the model make the voice or to record the words yourself with text to speech and then drive a face with them.

Sume's own rule, from the model docs, is plain: a shot where a person speaks on camera is Fabric with an accepted still plus TTS, because video models do not lip-sync to generated TTS or a later voice-over. So the choice is between an H3 clip whose audio you accept as it comes, and a TTS plus lip-sync pipeline whose words you control.

Price of a 10-second spoken line on Sume

All rates are the provider list times the 1.25 house margin, read from the repo's catalog on 2026-10-07. The TTS line assumes about 150 characters, which is $0.0071 and rounds up to $0.01.

10-second spoken line, Sume catalog read 2026-10-07
RouteRate10 secondsYou control the words
MiniMax H3, 768p, native audio$0.075 per second$0.75No, the model writes the sound
MiniMax H3 Max, 768p, native stereo audio$0.10 per second$1.00No
TTS + Fabric 480p$0.0475 per 1,000 characters + $0.10 per audio second$1.01Yes
TTS + Fabric 720p$0.0475 per 1,000 characters + $0.1875 per audio second$1.89Yes
TTS + H3 Max lip sync, 768p$0.0475 per 1,000 characters + $0.10 per audio second$1.01Yes
TTS + H3 Max lip sync, 1080p$0.0475 per 1,000 characters + $0.20 per audio second$2.01Yes

Limits that change the answer

  • H3 and H3 Max accept 5 to 15 seconds; H3 Max reads 480p, 768p or 1080p.
  • H3 Max lip sync takes audio of 5 to 14.8 seconds. Sume refuses both edges instead of clamping, so a 4-second line must be padded.
  • Fabric takes up to 300 seconds of audio, Sume-hosted, under 10 MB, so it suits long talking blocks.
  • A lip-sync clip's length follows the audio, so the price is audio seconds times the rate, rounded up per second.

When to pick which

Pick native audio when the sound is atmosphere: footsteps, a crowd, music under a product shot. The 768p H3 clip at $0.75 for ten seconds is the cheapest line in the table, and the sound is part of the price. Pick TTS plus a lip-sync model when a sentence has to be right, such as a price, a brand name or a legal line. The extra cost is about a cent of speech, and the lip-sync seconds cost what the video seconds would have.

Whichever route you use, listen to the result before it goes into a longer cut; the repo's tools describe the speech and sound as outputs to check, not guaranteed values.

Sources

Related posts

More in Models

All Models posts

Written by Sume