MiniMax H3 cache and sparse attention: what the open release has

MiniMax's H3 open release ships full-attention inference only; sparse attention comes later. About 13B parameters can be cached, not loaded.

4 min readSume
All posts

MiniMax says the initial H3 open-source release provides inference with full attention only; its sparse-attention implementation is not included and comes in a future update. It also says about 13B of the model's 33B parameters sit in AdaLN branches whose outputs can be precomputed and cached, so an inference-only deployment does not load them.

Both statements are on MiniMax's open-source announcement, read 2026-09-29. They describe the weights you download. On Sume you call a hosted model through the video API, where neither setting is a request field.

What does MiniMax say about the cache?

MiniMax describes the H3-Omni-Transformer as a 33B-parameter dense, single-stream Transformer, with approximately 13B parameters in AdaLN-related branches. Because the AdaLN modulation outputs can be precomputed and cached, those parameters do not need to be loaded for inference-only deployment. MiniMax says it releases the complete weights, fine-tuning included.

Is sparse attention in the open release?

From MiniMax's H3 open-source announcement, read 2026-09-29.
TopicWhat MiniMax's page says
Sparse attentionH3 natively supports sparse-attention training and inference
In the initial releaseInference with full attention only
Sparse implementationWill be released in a future update
AdaLN branchesAbout 13B of 33B parameters; outputs can be precomputed and cached

Does any of this change a Sume request?

No. Sume's video request has fields such as model, prompt, duration, resolution, aspect_ratio and generate_audio, and none of them controls attention or caching. Sume's docs say allowed_passthrough_parameters is empty for every model in v1, and a non-empty provider.options returns 400 unsupported_parameter. You cannot pass an engine setting through.

curl "https://api.sume.com/v1/videos/models" \
  -H "Authorization: Bearer $SUME_API_KEY"

Should I run the open weights myself instead?

That is a hosting decision, not a quality one. MiniMax's page also says the 2K regeneration module is not yet open-sourced, so the open release alone does not reproduce the 2K output it describes. Open-source video model or API? sets out the questions to ask, and MiniMax H3 in ComfyUI or through a hosted API covers the local route.

Does the cache mean H3 fits my GPU?

MiniMax's announcement does not say: it states that the cached parameters need not be loaded, but the page gives no memory or GPU requirement. Read the model card and the deployment notes for your hardware before you plan around it. On a hosted call there is no GPU to size.

What does a hosted call give up?

Control over the engine. On Sume you choose the model id, length, resolution, aspect ratio and references the catalog lists, and you get a job back. You cannot change the attention path or load your own fine-tune, and Sume's docs do not describe either as available. Fine-tuning is what the open weights are for, per MiniMax.

Sources

Related posts

More in Models

All Models posts

Written by Sume