MiniMax H3 prompting guide: the h3-prompt-writing skill
MiniMax ships nine H3 prompting skills in its repo; one is portable via npx skills add. What is public, what is Hub-only, and how to use the output on Sume.

MiniMax's H3 repository includes nine prompting skills, and one of them is portable: h3-prompt-writing, which the README installs with npx skills add https://github.com/MiniMax-AI/MiniMax-H3 --skill h3-prompt-writing. The README says the other eight style-specific video generation skills are exclusive to MiniMax Hub, so they are not in the repository.
What is public and what is not?
I did not read the contents of the skill, so this post does not summarize what it recommends. Read it before you rely on it.
| Resource | Availability |
|---|---|
| h3-prompt-writing skill | portable, installed with npx skills add |
| Eight style-specific video skills | MiniMax Hub only |
| Prompt guides for Base (FL2VA) and Reference (Ref2VA) modes | linked from the integration index |
Can the output be used on a hosted API?
A prompt is text. On Sume, minimax-h3 takes the same kind of string in the prompt field of POST /v1/videos, so a prompt written with the skill can be pasted in. What the skill cannot do is change the hosted row's limits: 5 to 15 seconds and 480p or 768p. Whether the row accepts seed is listed in its seed flag on GET /v1/videos/models.
If the guide you follow separates a first-and-last-frame mode from a reference mode, match it on Sume: frame_images with first_frame or last_frame for the first, input_references for the second. If you send both, frame_images decides and the request is treated as image-to-video.
How do you keep prompts portable?
Keep the scene in plain text and keep route-specific settings out of it. Duration, resolution, aspect ratio and references are separate fields on Sume, so a prompt that says "15 seconds, 2K" in the text does not set either. Put those in the request instead. That also lets you reuse the same prompt on a local run, where the same settings live in command flags.
Store prompts with the settings that produced the clip. When a result is good, you can reproduce it; when it is bad, you can change one thing at a time.
What should you test first?
Pick one scene, write it three ways, and render each at the lowest resolution that shows the idea. Keep the winner and then raise the resolution. The earlier posts on negative direction and multi-shot timecodes cover two prompt patterns on Sume. Request fields are in the video docs.
Sources
Related posts
More in Models
- MiniMax H3 VRAM: what 8 GB, 24 GB and B200 setups actually are
MiniMax H3 VRAM requirements are a ladder: 8 GB with NF4 offloading, 12 to 16 GB with GGUF, 24 GB with INT8, 8 x B200 for the reference server. Which tier fits.
- Mistral Large 4 on a 12-turn storyboard: tokens are 4% of the bill
Mistral Large 4 lists $1.36 input and $4.18 output per million tokens. A 12-turn run that makes six Wan 3.0 clips spends 4.3% of its total on tokens.
- Nano Banana 2.1 0.5K: Google says unsupported, Sume lists 512
Google's docs say the 512px tier is not supported on Nano Banana 2.1, yet the Sume catalog lists 512. Which to trust, and how to check the row live.
- Nano Banana 2.1 16:9 output size: 1376x768 up to 5504x3072
What pixel size a 16:9 Nano Banana 2.1 image comes out at per tier, and the 15 px crop that turns the 2K frame into exact 1920x1080.
Written by Sume