Choose the music in a Gemini Omni clip by prompting the audio

Gemini Omni makes its own soundtrack. Google's guide shows how to steer it with a music style, a radio effect or a timed chorus; here is a cheap way to test.

4 min readSume
All posts

To choose the music in a Gemini Omni clip, describe it in the prompt. Google's Omni guide says the model by default tries to generate an appropriate audio track, and that describing the audio you want is especially important if you want music. You cannot upload a song to match: Sume's Video Router doc says gemini-omni-flash-1.1 has native synced audio always on and no reference_audio_urls.

So the prompt is the only steering wheel, and it is worth knowing the wording Google documents.

What music wording does the guide give?

Under "Prompting the audio" the guide gives three example lines, and the extension section adds a fourth that names a song section.

Audio prompt examples from Google's Gemini Omni guide, read 2026-10-03
Prompt wordingWhat it asks for
Include calm background musicA calm bed under the scene
The video has a high energy techno beatA driving electronic track
The audio is a low tinny radio broadcast in the background, playing a songA song heard through a small radio
At 5s the chorus starts in the background audioA section change at a stated time
The music continues into the chorus (extension prompt)Continuing a track across an extended clip

How do I time a change in the music?

The guide's Timing events section says you can prompt for things to happen at specific times in natural language, with no precise syntax needed, and gives "At 5s the chorus starts in the background audio" as an example. It also offers a timecode form, such as [0-3s] and [3-6s] ranges. On Sume the clip length is 3 to 10 seconds, so a timed change has to fit inside that: a drop at 5 seconds in a 6-second clip leaves one second of the new section.

Keep the music request short and put it near the end of the prompt, after the scene description. This is a style habit rather than a documented rule, so test it on your own subject.

How do I test three music styles cheaply?

Run the same scene three times at the smallest size. Sume's Video Router doc lists gemini-omni-flash-1.1 at $0.03 per second at 360p on the provider list, so a 3-second draft lists at $0.09 before Sume's 1.25 multiplier. Change only the music sentence, keep everything else identical, and compare the three results by ear. Give each request its own Idempotency-Key.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-music-techno-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "A cyclist crosses a bridge at dawn, one continuous shot. The video has a high energy techno beat.",
    "resolution": "360p",
    "duration": 3,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

What if none of the generated music fits?

Replace it rather than fight the prompt. Sume's Timeline 1.0 has an optional soundtrack field with url, gain_db, loop, fade_out_seconds and duck_db, so a bed you already own can sit under a spoken line. If you want to keep the model's sound and add a bed beneath it, extract the clip's audio with audio detach first, which returns a wav or mp3 for $0.01 per job.

One reminder from Google's notes: every Omni video carries SynthID watermarking, so check each platform's AI disclosure rules before you publish.

Sources

Related posts

More in Models

All Models posts

Written by Sume