Can Gemini Omni Flash speak my script? It takes no audio input on Sume
On Sume, Gemini Omni Flash 1.1 takes image and video references but no audio reference. It makes its own sound. For your script in a face, use avatar video.

No. On Sume, Gemini Omni Flash 1.1 accepts image and video references and no audio reference, so you cannot hand it a recording or a voiceover and have a face say it. It generates its own native synced audio from the prompt.
For a person saying your exact words, use Sume's avatar video, which takes a script. For a clip with ambient sound, Omni is the right tool.
Who takes what
The video generation docs list supported_input_references per model. A model accepts a reference type only if it appears in that list.
| Model | Video reference | Audio reference |
|---|---|---|
| Gemini Omni Flash 1.1 | yes | no |
| Wan 3.0 | yes | yes |
| Seedance 2.x | yes | yes |
| MiniMax H3 and H3 Max | yes | yes |
What each route is for
Omni: product motion, scenes, ambient sound and edits to existing video. It makes audio, but the audio is the model's, not yours.
Wan 3.0, Seedance 2.x or MiniMax H3: when you want to supply an audio reference with the prompt. Wan 3.0 is $0.1250 a second at 720p.
Avatar video: when the words must be yours. It is a 4 to 60 second job on a ready avatar, from $0.184 a second on the Standard tier.
Narration over silent footage
Narration over product footage is a different build. Generate the visuals, record or synthesise the voice, then lay the voice down as the Timeline spine. Timeline renders at $0.10 per output minute, and voice is its own charge. At $1.25 for ten seconds of 720p Omni, the clip is the cheap part.
Pick the route by what the viewer must see. A mouth means avatar video. Hands and product mean Omni with a voiceover.
What I have not verified
I have not tested whether an Omni clip prompted with dialogue shows accurate lip movement, and Sume's docs do not claim it. If a spoken line has to be exact and sync to a mouth, use the avatar route, and check the first render by ear and eye.
A voiceover on top of a silent Omni clip works for narration over product footage, because nobody expects lips. Build that as a Timeline job with your audio as the spine.
Rates are from the Sume price card on 2026-10-07.
Sources
Related posts
More in Models
- Same character in three AI video shots: Omni reference image
Keep one character across separate Gemini Omni clips by sending the same reference image with every request on Sume. Request shape, prompt tags, 3-shot cost.
- ChatGPT Image 2 vs 2.5 at medium quality, 16:9: price on Sume
At 16:9 and medium quality, ChatGPT Image 2 bills $0.0525 per image on Sume and ChatGPT Image 2.5 bills $0.0105: five times less. Table, batch math and a swap.
- Do low, medium and high mean the same on every Sume image model?
No. On Sume only ChatGPT Image 2, 2.5 and Ideogram 4.5 accept low, medium and high, and the tiers change price in different ways. Defaults and the tier ladders.
- Does GPT Image 1 still work on October 7, 2026? 16 days left
OpenAI's deprecations page lists the GPT Image 1 shutdown for 2026-10-23, 16 days after 2026-10-07. What still works, what Sume serves and what to move to.
Written by Sume