MiniMax H3 camera prompts: lens, movement, exposure wording
fal's H3 prompting guide says H3 reads film vocabulary: lens, rack focus, handheld, grain. Wording that works as a prompt and a Sume request that sends it.

MiniMax H3 reads film vocabulary directly, so write the camera into the prompt the way a director's note would: name the lens, the movement, how the light behaves and what the film stock looks like. That is the guidance in fal's MiniMax H3 prompting guide (read 2026-10-02), and on Sume the same prompt goes to minimax-h3 or minimax-h3-max unchanged.
This post is about the camera layer only. Negative direction and reference jobs are in the prompt guide post, and shot labels with timecodes are in the multi-shot post. Nothing here is a Sume feature: Sume passes the text through, and the vocabulary list comes from fal's guide.
Which camera words does fal's guide name?
The guide groups four kinds of film language that H3 reads. I list only the examples the guide itself gives, so none of these wordings are my own invention.
Treat the table as starting phrases. The guide does not publish a pass rate for any of them, and I did not test them, so expect to iterate on short 5 second clips at 768p before you spend on a 15 second one.
| Layer | Example wording from the guide |
|---|---|
| Lens choice | wide angle lens with strong perspective distortion |
| Movement | rack focus; push in quickly; handheld shake |
| Exposure behavior | backlit breathing; coarse noise in shadows |
| Stock character | fine grain, soft highlight halation |
How do I structure a prompt around the camera?
Say what the shot is, then how the camera sees it. For anything beyond one beat, the guide recommends timecoded blocks such as "[0-2 seconds] High-angle overhead shot", which the guide says keeps a long output from becoming a slideshow. Put the camera note inside each block rather than once at the top, so each beat has its own lens and movement.
Pair every camera instruction with something to keep stable. The guide's advice for edits is to pair each change with what must stay stable; the same habit works for camera notes: name the move, then say what must not drift, such as the subject's wardrobe or the room's layout.
What does the Sume request look like?
A plain text-to-video call. minimax-h3 takes 5 to 15 seconds at native 480p or 768p, and minimax-h3-max takes 5 to 15 seconds at 480p, 768p or 1080p, with native stereo audio on both, per Sume's Video generation docs. The camera notes live in prompt.
The Idempotency-Key header makes a retry return the original job instead of billing a second clip.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-camera-001" \
-d '{
"model": "minimax-h3-max",
"prompt": "[0-3 seconds] Wide angle lens with strong perspective distortion, handheld shake, a cook flips a pan, backlit breathing. [3-6 seconds] Push in quickly to the pan, rack focus to the steam, fine grain, soft highlight halation.",
"duration": 6,
"resolution": "768p",
"aspect_ratio": "16:9"
}'What does this not guarantee?
Camera words steer a generative model; they do not command a rig. Handheld shake will not match a specific lens or a specific shot you have in mind, and the guide shows examples, not a spec. If you need the exact framing of a real shot, use first and last frame images (how that works) rather than prose.
Finally, 1080p on minimax-h3-max is a latent refinement from native 768p, so a lens or grain note is judged at 768p first. Pick the camera look at 768p, and raise the resolution only once the shot works.
Sources
Related posts
More in Models
- MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.
- MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.
- Where are the lyrics in a Lyria 3.5 result? Gemini vs Sume
Google returns Lyria 3.5 lyrics and song structure as text beside the audio. Sume puts model-reported lyrics or a section map in result.lyrics when present.
- Nano Banana negative prompt: describe what you want instead
Google's Gemini image guide says to write semantic negative prompts: describe an empty street, not 'no cars'. Sume's image request has no negative field.
Written by Sume