Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.

Veda is a learned sparse-attention predictor for MiniMax H3 that keeps the most important 10 percent of attention work and skips the rest, which its card says makes attention up to 6.8 times faster and whole-clip generation up to 3.1 times faster. The gap between those two numbers is the useful part: it tells you attention was most of the render time, and that the remaining work does not speed up.
Facts here are from the Veda-Sparse card and the Turbo LoRA card, read 2026-10-03. The Veda page is a preview and sits on Hugging Face's trending text-to-video list.
What does Veda change?
The card calls Veda a learned sparse-attention method for video diffusion models. A small predictor picks which attention operations matter, about the top 10 percent, and the model skips the others. It is plug-and-play and the card says it works with LoRAs and fine-tuned variants. Speedups grow with video length, according to the card. It was trained and evaluated with the Turbo LoRA for 8-step generation, so the card implies the best results at that setting and possible degradation at other step counts.
The card was tested on an RTX PRO 6000 Blackwell and an RTX 4090. On a 24 GB card it suggests memory flags, --offload-blocks 50 and --mlp-chunk-rows 4096. Like the other derivatives, it inherits the MiniMax H3 Community License from its base.
| Item | Card value |
|---|---|
| Attention work kept | About the most important 10 percent |
| Attention speedup | Up to 6.8 times |
| End-to-end speedup | Up to 3.1 times |
| Tested on | RTX PRO 6000 Blackwell and RTX 4090 |
| 24 GB cards | May need offload-blocks 50 and mlp-chunk-rows 4096 |
| Trained and evaluated with | MiniMax-H3 Turbo LoRA, 8 steps |
| Licence | Inherits the MiniMax H3 Community License |
Why is 3.1 times less than 6.8 times?
Because only part of the render is attention. A short calculation, which is arithmetic on the card's two numbers and not a figure the card states: if attention is a fraction f of the run time and becomes 6.8 times faster, the whole run speeds up by 1 divided by (1 minus f plus f over 6.8). Setting that equal to 3.1 gives f of roughly 0.79, so about four fifths of the baseline time was attention in the authors' setup.
That fraction is the ceiling on what any attention trick can do. Past this point the text encoder, the audio path, the decode and the non-attention layers dominate, and a faster attention kernel cannot help them. It also explains the card's remark that longer videos gain more: attention cost grows faster with length than the rest.
How does it stack with the Turbo LoRA?
The Turbo LoRA cuts the number of steps; Veda cuts the cost of each step. In principle the two multiply, and the card was trained against the 8-step Turbo setup, so it tells you the combination it targets. Neither card states the combined clip time on a named card, so measure it.
Run four timings on one prompt: base model, Turbo only, Veda only if the card's setup allows it, and both. Record end-to-end seconds, not step time, and judge motion and audio quality blind. A speedup that costs the audio sync or fine detail is not a gain for work you deliver.
Is any of this on a hosted route?
Not as options you can request. Sume's Video generation docs list minimax-h3 and minimax-h3-max with durations, resolutions and reference types; the request has no attention, step or adapter fields. You submit a job and read the result through Jobs and results.
So a preview like Veda is a local-iteration tool today. Use it if you own the card and generate many drafts; for finished clips with no setup time, the hosted route is the simpler choice, and the hosted price is in the model's pricing_skus.
Sources
Related posts
More in Models
- 9:16 vertical AI video: which Sume models take an aspect ratio
Seedance, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash take 9:16; Grok Imagine, Genjutsu and H3 Max Recast take no aspect ratio. A per-model table.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
Written by Sume