MiniMax H3 camera prompts: lens, movement, exposure wording

fal's H3 prompting guide says H3 reads film vocabulary: lens, rack focus, handheld, grain. Wording that works as a prompt and a Sume request that sends it.

5 min readSume
All posts

MiniMax H3 reads film vocabulary directly, so write the camera into the prompt the way a director's note would: name the lens, the movement, how the light behaves and what the film stock looks like. That is the guidance in fal's MiniMax H3 prompting guide (read 2026-10-02), and on Sume the same prompt goes to minimax-h3 or minimax-h3-max unchanged.

This post is about the camera layer only. Negative direction and reference jobs are in the prompt guide post, and shot labels with timecodes are in the multi-shot post. Nothing here is a Sume feature: Sume passes the text through, and the vocabulary list comes from fal's guide.

Which camera words does fal's guide name?

The guide groups four kinds of film language that H3 reads. I list only the examples the guide itself gives, so none of these wordings are my own invention.

Treat the table as starting phrases. The guide does not publish a pass rate for any of them, and I did not test them, so expect to iterate on short 5 second clips at 768p before you spend on a 15 second one.

Camera vocabulary from fal's H3 prompting guide (read 2026-10-02)
LayerExample wording from the guide
Lens choicewide angle lens with strong perspective distortion
Movementrack focus; push in quickly; handheld shake
Exposure behaviorbacklit breathing; coarse noise in shadows
Stock characterfine grain, soft highlight halation

How do I structure a prompt around the camera?

Say what the shot is, then how the camera sees it. For anything beyond one beat, the guide recommends timecoded blocks such as "[0-2 seconds] High-angle overhead shot", which the guide says keeps a long output from becoming a slideshow. Put the camera note inside each block rather than once at the top, so each beat has its own lens and movement.

Pair every camera instruction with something to keep stable. The guide's advice for edits is to pair each change with what must stay stable; the same habit works for camera notes: name the move, then say what must not drift, such as the subject's wardrobe or the room's layout.

What does the Sume request look like?

A plain text-to-video call. minimax-h3 takes 5 to 15 seconds at native 480p or 768p, and minimax-h3-max takes 5 to 15 seconds at 480p, 768p or 1080p, with native stereo audio on both, per Sume's Video generation docs. The camera notes live in prompt.

The Idempotency-Key header makes a retry return the original job instead of billing a second clip.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-camera-001" \
  -d '{
    "model": "minimax-h3-max",
    "prompt": "[0-3 seconds] Wide angle lens with strong perspective distortion, handheld shake, a cook flips a pan, backlit breathing. [3-6 seconds] Push in quickly to the pan, rack focus to the steam, fine grain, soft highlight halation.",
    "duration": 6,
    "resolution": "768p",
    "aspect_ratio": "16:9"
  }'

What does this not guarantee?

Camera words steer a generative model; they do not command a rig. Handheld shake will not match a specific lens or a specific shot you have in mind, and the guide shows examples, not a spec. If you need the exact framing of a real shot, use first and last frame images (how that works) rather than prose.

Finally, 1080p on minimax-h3-max is a latent refinement from native 768p, so a lens or grain note is judged at 768p first. Pick the camera look at 768p, and raise the resolution only once the shot works.

Sources

Related posts

More in Models

All Models posts

Written by Sume