Gemini Omni 'no dialogue' prompt: steer the audio on Sume

Omni clips always carry sound. Google suggests 'no dialogue' and 'no extra sound effects' in the prompt. How to use that on Sume and what a test costs.

4 min readSume
All posts

To keep a Gemini Omni clip free of speech, say so in the prompt: Google's Omni guide suggests negative phrasing such as 'no dialogue' and 'no extra sound effects' (read 2026-10-07). On Sume that wording is your only audio control, because the model always makes native synced audio and generate_audio: false is rejected with a 400.

Treat the phrase as a steer, not a switch. Test it at 360p first, listen, and only then pay for the resolution you ship.

What you can and cannot control

Sume's Video Router doc is explicit: gemini-omni-flash-1.1 has native synced audio with no toggle, and generate_audio: false is a 400 on this model. There is also no audio reference input, so you cannot hand it your own track. Everything audible is decided by the words in the prompt.

Audio levers for Omni (read 2026-10-07)
LeverGoogle's Omni guideOn Sume
Describe the sound you wantSay it explicitly, for example calm background musicWorks as prompt text
Negative wordingSuggests 'no dialogue' or 'no extra sound effects'Works as prompt text, not a guarantee
Turn audio offNot offered as a request field in the page readgenerate_audio: false returns 400
Audio reference inputCurrently unsupportedNo reference_audio_urls for this model

Wording that gives the steer a fair chance

Put the audio instruction in its own sentence at the end of the prompt, so it does not compete with the visual description. Name what you do want to hear instead of only banning things; a bare ban leaves the model free to fill the gap.

  • Positive first: 'Soft room tone and a distant street hum.'
  • Then the ban: 'No dialogue. No extra sound effects. No music.'
  • Avoid showing a mouth or a talking head if you want silence from speech; faces invite speech.
  • Keep one idea per sentence. Long run-on audio instructions are easier to ignore.

Send the request

This submits an 8-second 360p test to POST /v1/videos and polls the returned URL. The Idempotency-Key header makes a retried submit return the original job instead of a second charge.

import os, time, requests

H = {'Authorization': 'Bearer ' + os.environ['SUME_API_KEY']}
prompt = ('A potter shapes a bowl on a wheel, slow push-in, warm window light. '
          'Audio: soft room tone and the wet hum of the wheel. '
          'No dialogue. No extra sound effects. No music.')
body = {'model': 'gemini-omni-flash-1.1', 'prompt': prompt,
        'duration': 8, 'resolution': '360p', 'aspect_ratio': '16:9'}
r = requests.post('https://api.sume.com/v1/videos',
                  headers={**H, 'Idempotency-Key': 'potter-audio-test-1'},
                  json=body, timeout=60)
r.raise_for_status()
job = r.json()
while True:
    time.sleep(15)
    s = requests.get(job['polling_url'], headers=H, timeout=60).json()
    if s['status'] in ('completed', 'failed', 'cancelled'):
        break
print(s['status'], s.get('unsigned_urls'), s.get('error'))

What the test costs

Sume bills Omni at the provider list rate times 1.25 per second, rounded up to the cent per job. fal's list for Omni Flash 1.1 (recorded 2026-08-28) is $0.03, $0.10, $0.15 and $0.30 per second at 360p, 720p, 1080p and 4K, so an 8-second test is $0.30 at 360p and $1.00 at 720p. Three audio-wording tests at 360p cost $0.90, less than one 1080p final at $1.50.

If the steer fails and speech still appears, re-roll with a stronger positive sound description before you move up in resolution. If you need guaranteed silence, plan to remove the track after generation; the related post on dropping Omni audio covers the options Sume documents.

When the steer is not enough

If speech still appears after two re-rolls, stop spending at higher resolutions; the problem is the prompt, not the pixel count. Google's Omni page says English is fully supported and other languages are untested (read 2026-10-07), so keep the audio sentence in English even if the rest of your prompt is in another language. Sume's documented audio tools work on a finished clip: audio detach pulls the track out as a wav for $0.01 per job, and a timeline can lay a soundtrack bed under the picture. Neither is a one-click mute of the source, so decide before you generate whether you need a clean track or a clean picture.

Sources

Related posts

More in Models

All Models posts

Written by Sume