MiniMax H3 sound design prompts: direct the audio like the picture

fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.

5 min readSume
All posts

To control the sound of a MiniMax H3 clip, describe the audio in the prompt as specifically as you describe the picture: name the sonic elements, their weight and their placement, instead of writing "epic music". fal's MiniMax H3 prompting guide (read 2026-10-02) puts it as "direct the audio as deliberately as the picture", and gives the example of a deep sub-bass pulse, distant metallic resonance and restrained brass accents.

On Sume, minimax-h3 and minimax-h3-max always generate native stereo audio, so there is nothing to switch on and generate_audio cannot turn it off; see which models always add sound. The only lever you have over what you hear is the prompt, plus a voice reference clip if you want a specific voice.

What counts as specific audio direction?

The guide contrasts generic effect names with named elements. Use the second kind. A useful sound note answers three questions: what is making the sound, how it behaves over time, and where it sits against everything else.

Sume does not add anything to this. The prompt text goes through as written, and I have not seen a published pass rate for any of these phrasings, so treat each one as a draft to listen to.

Audio wording, generic versus specific, based on fal's H3 guide (read 2026-10-02)
Instead ofWrite something like
epic musicdeep sub-bass pulse, distant metallic resonance, restrained brass accents (the guide's own example)
footstepsslow boots on wet gravel, one step per second, close to the camera
ambient noiselow room tone, a distant ventilation hum, no music
a voiceone narrator, calm, close-mic, speaking the line in quotation marks

How do I handle spoken lines and a voice?

The guide says reference audio clips enable voice transfer, with the model matching vocal timing and emotional delivery to new dialogue, and that dialogue can be edited by giving both the original line and the new one. On Sume that route is a reference audio clip on H3 or H3 Max, covered in the voice reference post. Put the exact words in quotation marks so they are not paraphrased.

Keep spoken lines short for a 5 second clip. Sume's H3 rows run 5 to 15 seconds, and a long speech in a short clip is more likely to rush or truncate; that is my caution from the length rule, not a figure from the vendor.

What does the request look like?

A text-to-video call on the H3 Max row. Audio direction sits at the end of the prompt, after the picture, so it is easy to adjust without touching the shot.

Sume bills by output second at list times 1.25 per the Video Router docs, so test the sound on a 5 second clip first.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-sound-001" \
  -d '{
    "model": "minimax-h3-max",
    "prompt": "A lighthouse keeper climbs a spiral stair at night, slow push in. Audio: deep sub-bass pulse, distant metallic resonance, boots on iron steps, no music, no voice.",
    "duration": 5,
    "resolution": "768p",
    "aspect_ratio": "16:9"
  }'

What cannot you control?

You cannot ask Sume for a clip without sound from these two rows. If you want to replace the soundtrack, lay your own audio under the clip with the audio spine in Timeline.

Also note what the guide does not claim: it does not give a stereo-width control or a loudness target. "Stereo" is what the model outputs, not a setting. For what the stereo track means in practice, see the native stereo audio post.

Sources

Related posts

More in Models

All Models posts

Written by Sume