Gemini Omni 'no dialogue' prompt: steer the audio on Sume
Omni clips always carry sound. Google suggests 'no dialogue' and 'no extra sound effects' in the prompt. How to use that on Sume and what a test costs.

To keep a Gemini Omni clip free of speech, say so in the prompt: Google's Omni guide suggests negative phrasing such as 'no dialogue' and 'no extra sound effects' (read 2026-10-07). On Sume that wording is your only audio control, because the model always makes native synced audio and generate_audio: false is rejected with a 400.
Treat the phrase as a steer, not a switch. Test it at 360p first, listen, and only then pay for the resolution you ship.
What you can and cannot control
Sume's Video Router doc is explicit: gemini-omni-flash-1.1 has native synced audio with no toggle, and generate_audio: false is a 400 on this model. There is also no audio reference input, so you cannot hand it your own track. Everything audible is decided by the words in the prompt.
| Lever | Google's Omni guide | On Sume |
|---|---|---|
| Describe the sound you want | Say it explicitly, for example calm background music | Works as prompt text |
| Negative wording | Suggests 'no dialogue' or 'no extra sound effects' | Works as prompt text, not a guarantee |
| Turn audio off | Not offered as a request field in the page read | generate_audio: false returns 400 |
| Audio reference input | Currently unsupported | No reference_audio_urls for this model |
Wording that gives the steer a fair chance
Put the audio instruction in its own sentence at the end of the prompt, so it does not compete with the visual description. Name what you do want to hear instead of only banning things; a bare ban leaves the model free to fill the gap.
- Positive first: 'Soft room tone and a distant street hum.'
- Then the ban: 'No dialogue. No extra sound effects. No music.'
- Avoid showing a mouth or a talking head if you want silence from speech; faces invite speech.
- Keep one idea per sentence. Long run-on audio instructions are easier to ignore.
Send the request
This submits an 8-second 360p test to POST /v1/videos and polls the returned URL. The Idempotency-Key header makes a retried submit return the original job instead of a second charge.
import os, time, requests
H = {'Authorization': 'Bearer ' + os.environ['SUME_API_KEY']}
prompt = ('A potter shapes a bowl on a wheel, slow push-in, warm window light. '
'Audio: soft room tone and the wet hum of the wheel. '
'No dialogue. No extra sound effects. No music.')
body = {'model': 'gemini-omni-flash-1.1', 'prompt': prompt,
'duration': 8, 'resolution': '360p', 'aspect_ratio': '16:9'}
r = requests.post('https://api.sume.com/v1/videos',
headers={**H, 'Idempotency-Key': 'potter-audio-test-1'},
json=body, timeout=60)
r.raise_for_status()
job = r.json()
while True:
time.sleep(15)
s = requests.get(job['polling_url'], headers=H, timeout=60).json()
if s['status'] in ('completed', 'failed', 'cancelled'):
break
print(s['status'], s.get('unsigned_urls'), s.get('error'))What the test costs
Sume bills Omni at the provider list rate times 1.25 per second, rounded up to the cent per job. fal's list for Omni Flash 1.1 (recorded 2026-08-28) is $0.03, $0.10, $0.15 and $0.30 per second at 360p, 720p, 1080p and 4K, so an 8-second test is $0.30 at 360p and $1.00 at 720p. Three audio-wording tests at 360p cost $0.90, less than one 1080p final at $1.50.
If the steer fails and speech still appears, re-roll with a stronger positive sound description before you move up in resolution. If you need guaranteed silence, plan to remove the track after generation; the related post on dropping Omni audio covers the options Sume documents.
When the steer is not enough
If speech still appears after two re-rolls, stop spending at higher resolutions; the problem is the prompt, not the pixel count. Google's Omni page says English is fully supported and other languages are untested (read 2026-10-07), so keep the audio sentence in English even if the rest of your prompt is in another language. Sume's documented audio tools work on a finished clip: audio detach pulls the track out as a wav for $0.01 per job, and a timeline can lay a soundtrack bed under the picture. Neither is a one-click mute of the source, so decide before you generate whether you need a clean track or a clean picture.
Sources
Related posts
More in Models
- Gemini Omni sound effects prompt: name each sound in 8 seconds
How to prompt footsteps, a door and breaking glass in a Gemini Omni clip: Google says describe audio explicitly. Sume request, timing and cost per take.
- Qwen Image 2.1 sizes (2752x1536, 2400x1792) vs Sume Qwen ratios
The Qwen-Image-2.1 card lists 2048x2048, 2400x1792 and 2752x1536. Sume qwen-image takes 13 aspect_ratio values instead of pixels. How they line up.
- Qwen Image 2.1 Pro API: the model card lists no Pro variant
Searching for a Qwen Image 2.1 Pro API? The 2.1 model card names no Pro variant and no hosted endpoint. Here is what Sume lists for Qwen today.
- Qwen Image 2.1 Pro: what exists, and which Qwen ids Sume lists
People search for Qwen Image 2.1 Pro. Its model card lists no Pro variant; Sume's Image API lists qwen/qwen-image and qwen/qwen-image-max.
Written by Sume