Can Gemini Omni Flash speak my script? It takes no audio input on Sume

On Sume, Gemini Omni Flash 1.1 takes image and video references but no audio reference. It makes its own sound. For your script in a face, use avatar video.

4 min readSume
All posts

No. On Sume, Gemini Omni Flash 1.1 accepts image and video references and no audio reference, so you cannot hand it a recording or a voiceover and have a face say it. It generates its own native synced audio from the prompt.

For a person saying your exact words, use Sume's avatar video, which takes a script. For a clip with ambient sound, Omni is the right tool.

Who takes what

The video generation docs list supported_input_references per model. A model accepts a reference type only if it appears in that list.

Audio and video reference inputs by Sume video model, read 2026-10-07
ModelVideo referenceAudio reference
Gemini Omni Flash 1.1yesno
Wan 3.0yesyes
Seedance 2.xyesyes
MiniMax H3 and H3 Maxyesyes

What each route is for

Omni: product motion, scenes, ambient sound and edits to existing video. It makes audio, but the audio is the model's, not yours.

Wan 3.0, Seedance 2.x or MiniMax H3: when you want to supply an audio reference with the prompt. Wan 3.0 is $0.1250 a second at 720p.

Avatar video: when the words must be yours. It is a 4 to 60 second job on a ready avatar, from $0.184 a second on the Standard tier.

Narration over silent footage

Narration over product footage is a different build. Generate the visuals, record or synthesise the voice, then lay the voice down as the Timeline spine. Timeline renders at $0.10 per output minute, and voice is its own charge. At $1.25 for ten seconds of 720p Omni, the clip is the cheap part.

Pick the route by what the viewer must see. A mouth means avatar video. Hands and product mean Omni with a voiceover.

What I have not verified

I have not tested whether an Omni clip prompted with dialogue shows accurate lip movement, and Sume's docs do not claim it. If a spoken line has to be exact and sync to a mouth, use the avatar route, and check the first render by ear and eye.

A voiceover on top of a silent Omni clip works for narration over product footage, because nobody expects lips. Build that as a Timeline job with your audio as the spine.

Rates are from the Sume price card on 2026-10-07.

Sources

Related posts

More in Models

All Models posts

Written by Sume