AI video audio you cannot switch off: Omni 1.1, H3 and H3 Max on Sume
Gemini Omni Flash 1.1 rejects generate_audio false; MiniMax H3 and H3 Max have no toggle. To publish a clip with your own sound, drop the audio with video trim.

Three Sume video rows always make sound: Gemini Omni Flash 1.1 (which sume/auto uses), MiniMax H3 and MiniMax H3 Max. Omni 1.1 returns a 400 if you send generate_audio: false, and the two MiniMax rows have no toggle at all. If your Short or ad needs your own music or voice, remove the model's track after generation with Sume's video trim, using audio: "drop".
Which rows allow it
From the Video Router doc:
| Model | Native audio | Can you turn it off at generation? |
|---|---|---|
| gemini-omni-flash-1.1 (and sume/auto) | Always on, synced | No: generate_audio false fails with 400 |
| minimax-h3 | Native stereo | No toggle |
| minimax-h3-max | Native stereo | No toggle |
| Other rows | Per catalog generate_audio field | Check GET /v1/videos/models |
The fix: drop it after
Video trim has an audio field with keep (the default) and drop. A trim over the full length with drop returns a silent MP4. Then lay your voice track and music in Timeline, where the sound comes from an audio spine you supply plus an optional soundtrack bed. The trim limits are a source of up to 1800 seconds and an output of 0.2 to 900 seconds.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"start": 0,
"duration": 8,
"precision": "exact",
"audio": "drop"
}When to keep the sound
Dropping the track is not always right. A clip with synced effects, such as a pour or a footstep, can be more convincing with its own audio than with a music bed laid over silence. For a talking scene with generated speech, the sync is the whole point.
Decide per clip. For a product montage with a voiceover, drop the audio and build the sound in Timeline from your own spine. For a single hero shot that you publish as is, keep it. The two choices cost different amounts of work, and only the first adds a job.
Keep one more consideration in view. If you publish generated sound, read the disclosure and labelling rules of the platform you publish on, because audio can be part of what a label covers. That is not a technical limit, and the pages cited in this post do not address it.
Why bother
Platform pages list technical limits, not sound rules, so this is not about a spec. It is about control. A generated soundtrack may not match your brand or your voiceover, and music rights are your call, not the model's. The disclosure checklist is the better place to check what each platform wants labelled.
The cost is one more job, and exact trim re-encodes the picture. If that matters, mute in your editor instead.
Sources
Related posts
More in Developers
- AI video generator for business: build a request form from the API
Use GET /v1/formats and the io.input_kind field to build an internal video request form, grouped by what each Sume Format needs.
- AI video generator for business: no approval step over the API
Format runs started through the Sume API are unattended: approval gates are pre-granted. What that means for review, and the unattended_blocked failure.
- arq worker that polls an AI video job with defer_by in Python
An arq task reads Sume's job status once and enqueues itself again with _defer_by from next_poll_after_seconds, giving asyncio polling without a sleep loop.
- asyncio Semaphore size for Sume image batches: accepted capacity
Size the semaphore to what Sume accepts, concurrency plus queue: Free 6, Pro 24, Startup 48, Scale 120. A fake-submit test proves the peak never exceeds it.
Written by Sume