MiniMax-H3 Turbo LoRA: 4 to 8 steps, Apache 2.0, base licence
The MiniMax-H3 Turbo LoRA cuts sampling to 4 to 8 steps and lists Apache 2.0, but the 33B base keeps its own terms. What it needs and the hosted route.

The MiniMax-H3 Turbo LoRA is a small adapter, listed as Apache 2.0, that cuts H3 sampling from about 20 steps to as few as 4, which its card puts at roughly 5 times faster, while still producing video with synchronized stereo audio. It speeds up a model you already run; it does not shrink the base model, so you still need a large GPU. If you only want finished clips, Sume's hosted minimax-h3 runs the full model per job, with no step count to set.
Facts below are from the Turbo LoRA card, read 2026-10-03, where the repository appears on Hugging Face's trending text-to-video list. A related sparse-attention card is cited where it adds a licence fact.
What does the Turbo LoRA change?
The card describes a LoRA adapter for joint video and synchronized stereo audio generation with MiniMax-H3. It lowers sampling from roughly 20 steps to as few as 4. The card recommends 4 to 8 steps, says 6 to 8 looks noticeably better than 4, and says going past 8 gives no benefit and risks artifacts. Keep the LoRA strength at 1.0 and the scheduler on "simple". The recommended checkpoint is about 744 MB.
One more practical detail from the card: the 124 to 362 frame range at 24 fps maps to the same 5 to 15 second window Sume lists for the hosted model, so a prompt you tune locally at a given length transfers to a hosted job without changing its duration. Frame counts are not whole seconds, though, so round to a whole-second duration when you move a clip across.
| Item | Card value |
|---|---|
| Steps | About 20 down to 4; recommended 4 to 8 |
| Speed claim | Roughly 5 times faster sampling |
| Duration | 124 to 362 frames at 24 fps (5 to 15 seconds) |
| Resolution | Multiples of 32; short edge typically 768 |
| Audio | 32 kHz stereo |
| Base model size | About 33B parameters; 80 GB GPU recommended for the largest resolutions |
| Low-memory options | ComfyUI low_vram mode; a script flag trades about 13 GB of VRAM for CPU RAM |
| Licence listed | Apache 2.0 |
Does Apache 2.0 on the LoRA make the video free to sell?
Not by itself. The adapter is one file; the 33B base it modifies has its own terms. Cards for H3 derivatives, such as the Veda sparse-attention card, say they inherit the MiniMax H3 Community License from the base. Our earlier posts cover that licence: a 20 million revenue threshold and territory limits, among other conditions.
So read the base licence for what you deliver, and do not infer it from the badge on an adapter.
What does Sume's hosted minimax-h3 give you instead?
In Sume's Video generation docs, minimax-h3 accepts 5 to 15 seconds at native 480p or 768p (768p is first-class, not 720p), and minimax-h3-max is the faster 768p variant at 480p, 768p and 1080p for 5 to 15 seconds with native stereo audio. Billing is the provider list price times 1.25. The request fields in the docs are duration, resolution, aspect_ratio and the reference and frame fields; there is no step count or LoRA field, so a Turbo-style speedup is not something you can request.
That is the trade. Sume handles the GPU and the job lifecycle (Jobs and results); you give up control of steps and adapters.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-draft-001" \
-d '{
"model": "minimax-h3",
"prompt": "A barista pours oat milk into a flat white, steam rising, café ambience",
"resolution": "768p",
"duration": 8,
"aspect_ratio": "16:9",
"mode": "async"
}'When is the Turbo LoRA worth the setup?
It pays off when you generate many clips on a card you already own and iterate on prompts, because 4 to 8 steps changes how many drafts you get per hour. It does not pay off for a handful of finished clips, where the setup time exceeds the render time you would save.
Before adopting it, run the same prompt at 4, 6 and 8 steps and compare motion and audio, since the card itself says 6 to 8 is noticeably better than 4. If the draft quality is enough, keep Turbo for ideation and send only the winning prompts to a hosted job for the final render.
Sources
Related posts
More in Models
- MiniMax-Music3 weights: 5-minute songs, 8 GB, vs Sume Music Router
MiniMax-Music3 is open-weights: 5-minute 32 kHz songs, 8 GB with streaming, a $20M licence rule. What it needs, and what Sume's Music Router does instead.
- Nano Banana Pro interleaved text and images vs Sume image output
Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.
- OmniVoice: 600+ languages, CC-BY-NC weights, hosted TTS instead
OmniVoice covers 600+ languages in a 0.6B model, but its weights are CC-BY-NC. What the card says, what it omits, and where a hosted TTS job fits.
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
Written by Sume