Gemini Omni extend: no new dialogue on an uploaded talking clip
Google says you can't extend an uploaded Omni clip where someone talks to add dialogue; extension only appends, to clips up to 10 s. Sume lists no extend mode.

Google's Gemini Omni Flash guide says you cannot extend an uploaded video in which someone is talking in order to add more dialogue. Characters can stay silent in the extension, or you can extend a video the model generated by using multi-turn extension with previous_interaction_id. Sume's docs list four Omni modes (text, image, reference and edit) and no extend mode, so on Sume this limit has no counterpart to hit.
Facts here are from Google's Gemini API: Generate and edit videos with Gemini Omni Flash and Sume's Video Router, read 2026-10-02.
What are Google's other extension limits?
Google's limitations list says input videos for editing and extension must be 10 seconds or less when uploaded, and extension only appends to the end: prepending or extending the middle of a clip is not supported. It also says extending uploaded videos is not available in the European Economic Area, Switzerland and the United Kingdom, though extending model-generated videos is.
The guide adds that extension works in 10-second steps up to a total of 40 seconds, using the last 10 seconds of the original as context, and that some of the final frames of your input are edited to make the transition seamless.
| Case | What the guide says |
|---|---|
| Uploaded clip, silent characters | Extension allowed |
| Uploaded clip, someone talking, new dialogue wanted | Not supported |
| Model-generated clip, multi-turn | Supported with previous_interaction_id |
| Input length when uploading | 10 seconds or less |
| Where the new part goes | End of the video only |
Why does the talking case fail?
Google does not give a reason on the page, so do not guess one. What the page does give is a workaround: carry the scene forward in the same conversation, instead of re-uploading a finished clip.
What does Sume offer instead?
The Video Router page lists video_to_video as an edit: you send video_url, the prompt describes the change, and aspect_ratio and duration are not sent to the provider. A reference_video_urls entry conditions a fresh generation instead, with at most 3 clips of at most 3 seconds each, and the docs say Omni takes no reference audio.
Neither is an extension of the original timeline. If your plan is a longer talking scene, write the dialogue into the first generation's prompt and keep clips within the 3 to 10 second range Sume documents for Omni.
What should I decide before generating?
A few choices are cheaper to make up front.
- Decide on spoken lines before the first clip, not after.
- If you need Google's multi-turn extension, call Google's API directly; Sume's docs do not describe it.
- Check whether your users sit in the EEA, Switzerland or the UK before relying on uploaded-clip extension.
Sources
Related posts
More in Models
- Gemini Omni with several videos: 3 references, no cross-video use
Gemini Omni takes up to 3 reference clips of 3 s each, yet Google warns that reasoning across several videos may degrade output. Sume: one video_url source.
- Gemini Omni REST: output_video is SDK-only, read the steps array
Calling Gemini Omni over REST? interaction.output_video is SDK-only. Read the base64 video from the model_output step; on Sume you get a media URL instead.
- Gemini Omni [# Sources] and [# References] tags vs Sume fields
Gemini Omni binds media to roles with tags like <FIRST_FRAME> and [# Sources ...]. Sume uses request fields instead: image_url, end_image_url, reference lists.
- Gemini Omni and Veo in Korean: only English is fully supported
Google's Omni and Veo pages say English is fully supported; other languages aren't evaluated. For Korean prompts, describe in English and quote on-screen text.
Written by Sume