Gemini Omni sound effects prompt: name each sound in 8 seconds
How to prompt footsteps, a door and breaking glass in a Gemini Omni clip: Google says describe audio explicitly. Sume request, timing and cost per take.

To get specific sound effects from Gemini Omni, name each sound in the prompt and tie it to the action that makes it. Google's Omni guide says to describe the desired audio explicitly, and its Veo 3.1 guide says the same for effects and ambience (both read 2026-10-07). On Sume the model always produces native synced audio, so the prompt is where sound design happens.
A 10-second ceiling and a single audio track mean you list a few sounds, not a score.
Write the sound next to the action
A prompt that only says 'a thriller scene' leaves the sound to chance. A prompt that says 'heels on wet pavement, a metal door scraping open, a glass bottle shattering' gives the model three events to sync. Order them as they happen on screen.
- Use concrete nouns: 'heels on wet pavement', not 'footsteps sounds'.
- Say the surface or material: wood, gravel, tile, glass, metal.
- Add one ambience line for the room: 'low rain and a distant siren'.
- Keep it to three or four named sounds in 8 seconds; more tends to blur.
Combine with timecodes
Google's Omni page shows timecode syntax such as '[0-3s] person walks, [3-6s] they stop'. Put a sound in each range so the effect lands at a known second. The sample below is one 8-second request on POST /v1/videos.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: alley-sfx-1" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Single continuous shot in a rainy alley at night. [0-3s] a woman in heels walks toward camera, heels clicking on wet pavement. [3-6s] she pulls a metal door open, hinge scraping. [6-8s] a glass bottle falls and shatters. Ambience: steady rain, distant siren. No music.",
"duration": 8,
"resolution": "720p",
"aspect_ratio": "16:9"
}'Price per take and where Veo differs
Sume bills Omni at list times 1.25 per second, rounded up to the cent per job; fal's list for Omni Flash 1.1 (recorded 2026-08-28) is $0.03, $0.10, $0.15 and $0.30 per second at 360p, 720p, 1080p and 4K. An 8-second take is $0.30 at 360p and $1.00 at 720p. Sume does not list Veo 3.1, so the Veo numbers below are Google's own and are shown only for scale.
| Route | Per second | 8 seconds |
|---|---|---|
| Omni on Sume, 360p | $0.0375 | $0.30 |
| Omni on Sume, 720p | $0.125 | $1.00 |
| Veo 3.1 Fast 720p on Google | $0.10 | $0.80 |
| Veo 3.1 Lite 720p on Google | $0.05 | $0.40 |
Iterate cheaply
Sound is the part you only judge by listening, so draft at 360p. Change one sound per take; if you rewrite the whole prompt you cannot tell which phrase fixed the glass. Once the three effects land where you want them, re-run the same prompt at 720p or 1080p for the keeper. The model does not take a seed on Sume, so the final render is a new take, not a copy of the draft: listen again.
A worked example of ordering
Suppose the clip is a night alley. Write the visual beat first, then the sound that belongs to it, in the order a viewer would hear them: heels at the start, the hinge in the middle, the glass at the end. A sound listed out of order tends to land at the wrong second. If one effect is missing in a take, add its action to the visible description as well ('the bottle slips from her hand and shatters on the stones'), because the model ties sound to what it shows. If an effect is present but too loud, add 'quiet' or 'distant' to that one sound and leave the rest of the prompt alone, so the next take changes only that variable.
Sources
Related posts
More in Models
- Qwen Image 2.1 sizes (2752x1536, 2400x1792) vs Sume Qwen ratios
The Qwen-Image-2.1 card lists 2048x2048, 2400x1792 and 2752x1536. Sume qwen-image takes 13 aspect_ratio values instead of pixels. How they line up.
- Qwen Image 2.1 Pro API: the model card lists no Pro variant
Searching for a Qwen Image 2.1 Pro API? The 2.1 model card names no Pro variant and no hosted endpoint. Here is what Sume lists for Qwen today.
- Qwen Image 2.1 Pro: what exists, and which Qwen ids Sume lists
People search for Qwen Image 2.1 Pro. Its model card lists no Pro variant; Sume's Image API lists qwen/qwen-image and qwen/qwen-image-max.
- Remove an object from a photo without a mask: Ideogram 4.5 prompt edit
Remove a bin, cable or stray person from a photo with no mask: send it to ideogram/ideogram-v4.5 and name the object. $0.0375 to $0.275 per image on Sume.
Written by Sume