Choose the music in a Gemini Omni clip by prompting the audio
Gemini Omni makes its own soundtrack. Google's guide shows how to steer it with a music style, a radio effect or a timed chorus; here is a cheap way to test.

To choose the music in a Gemini Omni clip, describe it in the prompt. Google's Omni guide says the model by default tries to generate an appropriate audio track, and that describing the audio you want is especially important if you want music. You cannot upload a song to match: Sume's Video Router doc says gemini-omni-flash-1.1 has native synced audio always on and no reference_audio_urls.
So the prompt is the only steering wheel, and it is worth knowing the wording Google documents.
What music wording does the guide give?
Under "Prompting the audio" the guide gives three example lines, and the extension section adds a fourth that names a song section.
| Prompt wording | What it asks for |
|---|---|
| Include calm background music | A calm bed under the scene |
| The video has a high energy techno beat | A driving electronic track |
| The audio is a low tinny radio broadcast in the background, playing a song | A song heard through a small radio |
| At 5s the chorus starts in the background audio | A section change at a stated time |
| The music continues into the chorus (extension prompt) | Continuing a track across an extended clip |
How do I time a change in the music?
The guide's Timing events section says you can prompt for things to happen at specific times in natural language, with no precise syntax needed, and gives "At 5s the chorus starts in the background audio" as an example. It also offers a timecode form, such as [0-3s] and [3-6s] ranges. On Sume the clip length is 3 to 10 seconds, so a timed change has to fit inside that: a drop at 5 seconds in a 6-second clip leaves one second of the new section.
Keep the music request short and put it near the end of the prompt, after the scene description. This is a style habit rather than a documented rule, so test it on your own subject.
How do I test three music styles cheaply?
Run the same scene three times at the smallest size. Sume's Video Router doc lists gemini-omni-flash-1.1 at $0.03 per second at 360p on the provider list, so a 3-second draft lists at $0.09 before Sume's 1.25 multiplier. Change only the music sentence, keep everything else identical, and compare the three results by ear. Give each request its own Idempotency-Key.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-music-techno-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A cyclist crosses a bridge at dawn, one continuous shot. The video has a high energy techno beat.",
"resolution": "360p",
"duration": 3,
"aspect_ratio": "9:16",
"mode": "async"
}'What if none of the generated music fits?
Replace it rather than fight the prompt. Sume's Timeline 1.0 has an optional soundtrack field with url, gain_db, loop, fade_out_seconds and duck_db, so a bed you already own can sit under a spoken line. If you want to keep the model's sound and add a bed beneath it, extract the clip's audio with audio detach first, which returns a wav or mp3 for $0.01 per job.
One reminder from Google's notes: every Omni video carries SynthID watermarking, so check each platform's AI disclosure rules before you publish.
Sources
Related posts
More in Models
- Gemini Omni rapid-fire video: a new labelled item every second
Prompt Gemini Omni for a rapid-fire clip that shows a different item every second with a text label, then send it through Sume's video router in 9:16.
- AI video signs and plates garbled? Write the text in the Omni prompt
Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.
- German and Italian text to speech API: de and it on Sume TTS
German (de) and Italian (it) are in both Cartesia Sonic 3.6 and Sume's voice library. Send language de or it, reuse one voice, and watch the 409 language check.
- Google Pics API? Edit one object with Nano Banana on Sume
Google's Pics announcement describes an app, not an API. The closest call on Sume: Nano Banana reference edits, or a GPT Image 2.5 mask for one region.
Written by Sume