Prompting Omni Flash with sound: write the audio as its own line

Gemini Omni Flash 1.1 always generates audio. A prompt layout that separates picture from sound, three example prompts, and the request on Sume.

3 min readSume
All posts

Because Omni Flash always produces native synced audio, the prompt is also your sound brief. A simple layout is: one sentence of picture, one sentence of camera, one sentence that starts with "Sound:". Nothing in the docs requires that format; it is a habit that keeps audio from being left to chance.

Sume's Video Router page states the audio is always on and that generate_audio: false is rejected, and the model takes no audio input, so you cannot supply a voice track to follow.

Three prompts

Example wording, not tested outputs
UsePicture and cameraSound line
Cafe hookA barista pours latte art into a white cup. Close, slow push in.Sound: milk pouring, a spoon on porcelain, quiet cafe murmur.
Shoe dropWhite sneakers on wet pavement, low angle, static.Sound: rain on asphalt, one footstep, no music.
Tool demoA cordless drill drives a screw into oak. Side view.Sound: motor whine rising, then a short stop.

Request

The sound line goes into the same prompt string.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sound-test-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "A barista pours latte art into a white cup. Close, slow push in. Sound: milk pouring, a spoon on porcelain, quiet cafe murmur.",
    "duration": 5,
    "resolution": "360p",
    "aspect_ratio": "9:16"
  }'

Test cheaply

Audio is in the price, so a sound test costs the same as a silent one: $0.1875 for 5 seconds at 360p. Run three wordings at 360p ($0.5625) and keep the one you like before you spend on 1080p.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume