AI video with native stereo sound: MiniMax H3 audio through the API
MiniMax H3 generates native stereo sound in the same pass as the video. On Sume audio is always on, generate_audio false is refused, and audio steers.

MiniMax H3 generates its sound with the picture: MiniMax's announcement says H3 produces video with jointly generated audio and that all audio output is native stereo. On Sume both minimax-h3 and minimax-h3-max always return that audio, so you leave generate_audio out; sending false is refused.
Vendor facts are from MiniMax's announcement and model card and fal's prompting guide; Sume facts from the Video generation docs and request validation code, read 2026-09-29.
Can I turn the audio off?
Not on these ids. In current code a request with generate_audio: false for minimax-h3 or minimax-h3-max fails validation with a message to omit the field, and bitrate_mode is refused too. If you need a silent clip, generate it and mute or replace the track in your own editing step.
How do I direct the sound in the prompt?
Say it in the prompt. MiniMax's own prompt guide structures a prompt around a description plus an overall_soundscape and non_diegetic_music field, suggesting 1–4 English sentences for the soundscape and 1–3 for the music. It writes dialogue in <d>[Language] text</d> tags and gives speakers stable ids like (S1). fal's guide advises directing audio deliberately, for example naming small sounds such as "clothing movement" and "subtle room air".
| Element | Convention in the guide |
|---|---|
| Dialogue | <d>[Language] text</d>, kept verbatim, not translated |
| Speakers | Stable ids such as (S1), (S2) |
| Voiceover | The phrase "says in an off-screen voiceover" |
| Ambience | overall_soundscape, 1–4 English sentences |
| Score | non_diegetic_music, 1–3 English sentences |
Can I send my own audio?
Yes, as a reference: up to 3 audio clips, each 2–15 seconds and 15 seconds combined, but audio cannot be the only reference; pair it with an image or a video. See voice reference with audio clips.
What format is the audio?
MiniMax's model card states 32 kHz stereo audio at 24 FPS video. Sume's docs call the audio native stereo and do not restate the sample rate, so treat the card as the vendor's figure.
Sources
Related posts
More in Models
- MiniMax H3 open weights on Hugging Face: what is in the release
MiniMax H3 has a model card on Hugging Face: two checkpoints, a 33B-parameter Transformer, a community license and a 4-GPU serving example. What to check first.
- AI video editing with a text prompt: MiniMax H3 and Sume's edit path
MiniMax H3 is described as editing existing video from instructions. What the vendor says, what Sume's H3 ids accept, and the id with an edit field.
- Nano Banana 2 Lite API: what Google shipped and what Sume lists
Nano Banana 2 Lite is Google's gemini-3.1-flash-lite-image model. What Google says it does, and which Nano Banana models Sume's image API lists today.
- Nano Banana Pro vs Nano Banana 2: which id to send on Sume
Nano Banana Pro and Nano Banana 2 share tiers and 10 references on Sume; they differ in extra aspect ratios and in which tiers change the price.
Written by Sume