Gemini Omni extend: no new dialogue on an uploaded talking clip

Google says you can't extend an uploaded Omni clip where someone talks to add dialogue; extension only appends, to clips up to 10 s. Sume lists no extend mode.

5 min readSume
All posts

Google's Gemini Omni Flash guide says you cannot extend an uploaded video in which someone is talking in order to add more dialogue. Characters can stay silent in the extension, or you can extend a video the model generated by using multi-turn extension with previous_interaction_id. Sume's docs list four Omni modes (text, image, reference and edit) and no extend mode, so on Sume this limit has no counterpart to hit.

Facts here are from Google's Gemini API: Generate and edit videos with Gemini Omni Flash and Sume's Video Router, read 2026-10-02.

What are Google's other extension limits?

Google's limitations list says input videos for editing and extension must be 10 seconds or less when uploaded, and extension only appends to the end: prepending or extending the middle of a clip is not supported. It also says extending uploaded videos is not available in the European Economic Area, Switzerland and the United Kingdom, though extending model-generated videos is.

The guide adds that extension works in 10-second steps up to a total of 40 seconds, using the last 10 seconds of the original as context, and that some of the final frames of your input are edited to make the transition seamless.

Google's Gemini Omni Flash guide, read 2026-10-02.
CaseWhat the guide says
Uploaded clip, silent charactersExtension allowed
Uploaded clip, someone talking, new dialogue wantedNot supported
Model-generated clip, multi-turnSupported with previous_interaction_id
Input length when uploading10 seconds or less
Where the new part goesEnd of the video only

Why does the talking case fail?

Google does not give a reason on the page, so do not guess one. What the page does give is a workaround: carry the scene forward in the same conversation, instead of re-uploading a finished clip.

What does Sume offer instead?

The Video Router page lists video_to_video as an edit: you send video_url, the prompt describes the change, and aspect_ratio and duration are not sent to the provider. A reference_video_urls entry conditions a fresh generation instead, with at most 3 clips of at most 3 seconds each, and the docs say Omni takes no reference audio.

Neither is an extension of the original timeline. If your plan is a longer talking scene, write the dialogue into the first generation's prompt and keep clips within the 3 to 10 second range Sume documents for Omni.

What should I decide before generating?

A few choices are cheaper to make up front.

  • Decide on spoken lines before the first clip, not after.
  • If you need Google's multi-turn extension, call Google's API directly; Sume's docs do not describe it.
  • Check whether your users sit in the EEA, Switzerland or the UK before relying on uploaded-clip extension.

Sources

Related posts

More in Models

All Models posts

Written by Sume