Gemini Omni sound effects prompt: name each sound in 8 seconds

How to prompt footsteps, a door and breaking glass in a Gemini Omni clip: Google says describe audio explicitly. Sume request, timing and cost per take.

4 min readSume
All posts

To get specific sound effects from Gemini Omni, name each sound in the prompt and tie it to the action that makes it. Google's Omni guide says to describe the desired audio explicitly, and its Veo 3.1 guide says the same for effects and ambience (both read 2026-10-07). On Sume the model always produces native synced audio, so the prompt is where sound design happens.

A 10-second ceiling and a single audio track mean you list a few sounds, not a score.

Write the sound next to the action

A prompt that only says 'a thriller scene' leaves the sound to chance. A prompt that says 'heels on wet pavement, a metal door scraping open, a glass bottle shattering' gives the model three events to sync. Order them as they happen on screen.

  • Use concrete nouns: 'heels on wet pavement', not 'footsteps sounds'.
  • Say the surface or material: wood, gravel, tile, glass, metal.
  • Add one ambience line for the room: 'low rain and a distant siren'.
  • Keep it to three or four named sounds in 8 seconds; more tends to blur.

Combine with timecodes

Google's Omni page shows timecode syntax such as '[0-3s] person walks, [3-6s] they stop'. Put a sound in each range so the effect lands at a known second. The sample below is one 8-second request on POST /v1/videos.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: alley-sfx-1" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Single continuous shot in a rainy alley at night. [0-3s] a woman in heels walks toward camera, heels clicking on wet pavement. [3-6s] she pulls a metal door open, hinge scraping. [6-8s] a glass bottle falls and shatters. Ambience: steady rain, distant siren. No music.",
    "duration": 8,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

Price per take and where Veo differs

Sume bills Omni at list times 1.25 per second, rounded up to the cent per job; fal's list for Omni Flash 1.1 (recorded 2026-08-28) is $0.03, $0.10, $0.15 and $0.30 per second at 360p, 720p, 1080p and 4K. An 8-second take is $0.30 at 360p and $1.00 at 720p. Sume does not list Veo 3.1, so the Veo numbers below are Google's own and are shown only for scale.

Cost of one 8-second take with audio (read 2026-10-07)
RoutePer second8 seconds
Omni on Sume, 360p$0.0375$0.30
Omni on Sume, 720p$0.125$1.00
Veo 3.1 Fast 720p on Google$0.10$0.80
Veo 3.1 Lite 720p on Google$0.05$0.40

Iterate cheaply

Sound is the part you only judge by listening, so draft at 360p. Change one sound per take; if you rewrite the whole prompt you cannot tell which phrase fixed the glass. Once the three effects land where you want them, re-run the same prompt at 720p or 1080p for the keeper. The model does not take a seed on Sume, so the final render is a new take, not a copy of the draft: listen again.

A worked example of ordering

Suppose the clip is a night alley. Write the visual beat first, then the sound that belongs to it, in the order a viewer would hear them: heels at the start, the hinge in the middle, the glass at the end. A sound listed out of order tends to land at the wrong second. If one effect is missing in a take, add its action to the visible description as well ('the bottle slips from her hand and shatters on the stones'), because the model ties sound to what it shows. If an effect is present but too loud, add 'quiet' or 'distant' to that one sound and leave the rest of the prompt alone, so the next take changes only that variable.

Sources

Related posts

More in Models

All Models posts

Written by Sume