Gemini Omni reference video: is the audio used? No
Gemini Omni Flash accepts video references but not audio: Google says audio in a reference clip is ignored. Limits, Sume's fields, and models that take audio.

No. A reference video for Gemini Omni Flash brings its picture, not its sound. Google's Omni page says any audio in a video reference is ignored and that uploading audio references is unsupported, and Sume's docs say gemini-omni-flash-1.1 accepts video references but not audio.
Both statements were read on 2026-09-29: Google's Omni page and Sume's Video generation docs.
What are the reference limits?
Google says video references work best with likenesses and support at most 3 clips of up to 3 seconds each. Sume's Video Router docs match on video and add an image cap: up to 10 reference images and up to 3 reference videos, each 3 seconds or less.
| Input | Accepted | Limit |
|---|---|---|
| Reference images | Yes | Up to 10 |
| Reference videos | Yes, picture only | Up to 3, each 3 s or less |
| Reference audio | No | No reference_audio_urls field |
Then where does the sound come from?
From the model. Sume's docs say native audio is always on for gemini-omni-flash-1.1, and generate_audio: false is rejected. Describe the sound you want in the prompt instead of supplying a clip; Google's guide says the model tries to generate an appropriate track and that describing the audio matters.
Which Sume models take an audio reference?
The docs name the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max as honoring audio and video references. If a voice or a music bed has to drive the clip, pick one of those rather than Omni.
How do I send a video reference?
Through the Video Router, use reference_video_urls and address the clip as <VIDEO_REF_0> in the prompt; media is numbered from 0 in list order.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-ref-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "The person in <VIDEO_REF_0> plays the violin on a rooftop at dusk",
"reference_video_urls": ["https://example.com/ref.mp4"],
"mode": "async"
}'What if my reference clip has a voice I want to keep?
Then Omni will not carry it. Because Google says audio in a reference video is ignored, the voice, music or room tone in your clip never reaches the model, and the new clip gets a generated track. You can describe the voice in the prompt, but a described voice is not your recording.
The practical workaround is to split the job. Generate the picture with Omni, then add your own voice or music afterwards in an editing step. Or start from a model in the list above that honors audio references, so the sound is part of the generation.
Do reference images change the audio question?
No. Images carry no sound, so up to 10 of them steer the look of the clip while the audio stays generated. Mixing images with a short video reference is allowed by the limits above; the only input Omni lacks is audio. If you find a request rejected, check the field names against the Video Router docs before assuming the model cannot do it.
The catalog is the current source of truth. GET /v1/video-router/models returns each model's capabilities, and Sume's docs advise reading those rather than assuming one shape for every model.
Sources
Related posts
More in Models
- Does Gemini Omni video have a SynthID watermark?
Google says all Gemini Omni Flash videos include an invisible SynthID watermark. What that means for API clips, YouTube disclosure, and what Sume's docs say.
- Google Flow video length: 4 to 10 seconds by model
Google Flow makes 4, 6 and 8 second clips with Veo 3.1 and 4 to 10 seconds with Gemini Omni Flash 1.1. Length by model and mode, from Google's help page.
- Does Sume Auto use GPT Image 2.5? What sume/auto picks for images
Sume's docs say Auto image routing continues to use Flare, GPT Image 2.5. Sume never says which family ran on a call, so pin openai/gpt-image-2.5 if it matters.
- GPT Image 2.5 Flare vs Sunburst: what each model page says
OpenAI calls Flare its model for everyday image generation and Sunburst its most capable for generation and editing. What each page lists, and both ids on Sume.
Written by Sume