DMAD 4-step MiniMax H3: 50-step baseline vs Turbo LoRA
DMAD distills MiniMax H3 from 50 steps to 4. Why that 4 is not the Turbo LoRA's 4, what the card says, what it omits, and where hosted jobs fit.

DMAD is an adversarially distilled student of MiniMax H3 that cuts inference from 50 steps to 4 while keeping joint video and stereo audio, shipped as rank-128 LoRA adapters. Its "4 steps" is not comparable with the Turbo LoRA's "4 steps", because the two cards start from different baselines: 50 steps for DMAD, about 20 for the Turbo LoRA. Compare speedups only against a shared baseline you measured yourself.
Facts are from the DMAD card and the Turbo LoRA card, both read 2026-10-03; DMAD is listed on Hugging Face's trending text-to-video page.
What is DMAD, according to its card?
The card describes a distilled student model for fast visual generation through adversarial distillation. It enables 4-step MiniMax-H3 students for joint audio and video generation, producing 1344x768 video at 24 fps with native stereo audio. The adapters are rank-128 LoRAs applied across 312 modules of the H3 transformer, and two checkpoints are offered: the paper's EMA version at iteration 800 and a fully trained variant at iteration 1600.
The card does not state a GPU requirement or usage limits, and it says you must download separate model components, so you still need the base H3 stack.
Treat the 1344x768 figure as what the authors report for their evaluation, not as a size Sume's hosted route will return. Sume's catalog advertises the resolutions each model accepts in supported_resolutions, so read that field for the hosted model instead of copying a size from a local-run card.
| Item | DMAD | Turbo LoRA |
|---|---|---|
| Method | Adversarial distillation into a student | LoRA for faster sampling |
| Baseline steps | 50 | About 20 |
| Target steps | 4 | 4 to 8 recommended |
| Adapter | Rank-128 LoRA, 312 modules | LoRA, recommended file about 744 MB |
| Output on the card | 1344x768 at 24 fps, stereo audio | Short edge typically 768, 32 kHz stereo |
| Terms on the card | MiniMax H3 Community License, as derivatives | Apache 2.0 listed |
Why can't you compare '4 steps' across the two?
A step count only means something against the step count it replaced. Going from 50 to 4 removes 46 steps; going from about 20 to 4 removes 16. The two ratios, 50 to 4 and about 20 to 4, are not the same claim, and neither one tells you which setup renders a clip sooner on your card. Neither card publishes a clip time on a named GPU for this comparison.
What you can compare is your own measurement: same prompt, same resolution, same card, same length, timed from submit to a saved file, with the quality judged by someone who did not pick the settings. Keep the seed fixed when the runtime honors one.
What does the licence line say?
The DMAD card states that it is distributed under the MiniMax H3 Community License Agreement as model derivatives. The Turbo LoRA card lists Apache 2.0 for the adapter. The mismatch is a reminder that each card is the source for its own file, and that the base model's terms can follow you either way. Our posts on that licence cover its revenue threshold and its territory exclusions.
Read the licence against how you deliver the clip, and keep a copy of the version you accepted.
When is a hosted job the simpler route?
When you do not want to run the stack. Sume's Video generation docs list minimax-h3 for 5 to 15 seconds at native 480p or 768p and minimax-h3-max as the faster 768p variant, with native stereo audio. The request has no step count and no adapter field, so you cannot ask for a distilled student; you submit a prompt and read the finished job.
Use a student like DMAD when you generate many drafts locally and want faster iteration, then hand the final prompt to a hosted job when you want the full model's render without owning the GPU. If you do, pin the model id explicitly rather than sending sume/auto, so the final clip comes from the model you tested.
Sources
Related posts
More in Models
- End-frame AI video: which Sume models take a last frame
Seedance, Wan 3.0, Kling 3, MiniMax and Gemini Omni Flash accept a first and last frame; Grok Imagine and the swap rows do not. How to send both frames.
- Fastest Seedance or Kling model on Sume: latency labels
Sume labels seedance-2-fast and seedance-2-mini fast, seedance-2.5 and seedance-2 medium, and kling-3 medium-slow. What the labels mean and how to test.
- FLUX 3 Image 768sq and 1.5k tiers vs Sume's 512, 1K, 2K, 4K
FLUX 3 Image sells 768sq, 1k, 1.5k, 2k and 4k; Sume's Nano Banana tiers are 512, 1K, 2K, 4K. A Python tier translator that rounds up and checks the catalog.
- FLUX 3 Image open weights: release date and what to use now
FLUX 3 Image's open weights are reported as coming within weeks, with no date. What is known, what is not, and the hosted FLUX.2 route on Sume today.
Written by Sume