MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.

To control the sound of a MiniMax H3 clip, describe the audio in the prompt as specifically as you describe the picture: name the sonic elements, their weight and their placement, instead of writing "epic music". fal's MiniMax H3 prompting guide (read 2026-10-02) puts it as "direct the audio as deliberately as the picture", and gives the example of a deep sub-bass pulse, distant metallic resonance and restrained brass accents.
On Sume, minimax-h3 and minimax-h3-max always generate native stereo audio, so there is nothing to switch on and generate_audio cannot turn it off; see which models always add sound. The only lever you have over what you hear is the prompt, plus a voice reference clip if you want a specific voice.
What counts as specific audio direction?
The guide contrasts generic effect names with named elements. Use the second kind. A useful sound note answers three questions: what is making the sound, how it behaves over time, and where it sits against everything else.
Sume does not add anything to this. The prompt text goes through as written, and I have not seen a published pass rate for any of these phrasings, so treat each one as a draft to listen to.
| Instead of | Write something like |
|---|---|
| epic music | deep sub-bass pulse, distant metallic resonance, restrained brass accents (the guide's own example) |
| footsteps | slow boots on wet gravel, one step per second, close to the camera |
| ambient noise | low room tone, a distant ventilation hum, no music |
| a voice | one narrator, calm, close-mic, speaking the line in quotation marks |
How do I handle spoken lines and a voice?
The guide says reference audio clips enable voice transfer, with the model matching vocal timing and emotional delivery to new dialogue, and that dialogue can be edited by giving both the original line and the new one. On Sume that route is a reference audio clip on H3 or H3 Max, covered in the voice reference post. Put the exact words in quotation marks so they are not paraphrased.
Keep spoken lines short for a 5 second clip. Sume's H3 rows run 5 to 15 seconds, and a long speech in a short clip is more likely to rush or truncate; that is my caution from the length rule, not a figure from the vendor.
What does the request look like?
A text-to-video call on the H3 Max row. Audio direction sits at the end of the prompt, after the picture, so it is easy to adjust without touching the shot.
Sume bills by output second at list times 1.25 per the Video Router docs, so test the sound on a 5 second clip first.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-sound-001" \
-d '{
"model": "minimax-h3-max",
"prompt": "A lighthouse keeper climbs a spiral stair at night, slow push in. Audio: deep sub-bass pulse, distant metallic resonance, boots on iron steps, no music, no voice.",
"duration": 5,
"resolution": "768p",
"aspect_ratio": "16:9"
}'What cannot you control?
You cannot ask Sume for a clip without sound from these two rows. If you want to replace the soundtrack, lay your own audio under the clip with the audio spine in Timeline.
Also note what the guide does not claim: it does not give a stereo-width control or a loudness target. "Stereo" is what the model outputs, not a setting. For what the stereo track means in practice, see the native stereo audio post.
Sources
Related posts
More in Models
- Where are the lyrics in a Lyria 3.5 result? Gemini vs Sume
Google returns Lyria 3.5 lyrics and song structure as text beside the audio. Sume puts model-reported lyrics or a section map in result.lyrics when present.
- Nano Banana negative prompt: describe what you want instead
Google's Gemini image guide says to write semantic negative prompts: describe an empty street, not 'no cars'. Sume's image request has no negative field.
- Nano Banana prompt languages: Korean and Japanese, per Google
Google lists the languages Gemini image models work best in, including ko-KR and ja-JP. What that means for a Korean or Japanese prompt sent through Sume.
- Nano Banana text in images: write the copy first, then render
Google's tip for text in Gemini images: settle the wording first, then ask for the image. How to do that with one Sume request and a short checklist.
Written by Sume