Gemini Omni extension only appends: how to add a lead-in clip
Google's Omni docs say extension adds to the end of a clip only. To get a clip before your footage, generate it separately and join on a Sume timeline.

Can Gemini Omni extend a video backwards, or insert a scene in the middle? No. The Gemini API Omni docs say extension appends to the end only, with no prepend and no middle insert, and that uploaded input videos must be 10 seconds or shorter. To get a lead-in before existing footage, create the lead-in as its own clip and join the two.
What Google documents
Google's launch post says scene extension analyzes up to 10 seconds of prior context and extends in 10 second increments up to 40 seconds. That is a forward-only chain.
| Question | Answer on the vendor page |
|---|---|
| Extend at the end | Yes, up to 40 s total |
| Prepend before the start | No |
| Insert in the middle | No |
| Longest uploaded input | 10 s |
The Sume route
Sume's gemini-omni-flash-1.1 row does not expose extend or previous_interaction_id; its edit mode takes a video_url. So the lead-in plan is an assembly plan, not an extension.
- Pull the first frame of your footage with video-frames using
at[]. - Generate the lead-in with that frame as the last-frame input, so the lead-in ends on a frame that matches your footage.
- Join the lead-in and your clip on a timeline render.
Limits to plan for
Sume lists Omni at 3 to 10 seconds, 16:9 and 9:16 only, so the lead-in must match the aspect ratio of the footage. The first-frame match will not hide a change in lighting or motion, so keep the lead-in short and expect a visible cut. Generate two or three candidates and choose the one whose last frame sits closest to the footage.
A note on what Sume does not do
Sume does not run Veo or any model absent from its video catalog, and it does not expose Google's extend feature for Omni. If your workflow depends on the Gemini app's scene extension, that runs in Google's own products: Flow for AI Plus, Pro and Ultra subscribers, and the Gemini app, according to the launch post.
Sequence to run
Work backward from the footage you already have. First decide how long the lead-in should be; Omni on Sume produces 3 to 10 seconds, so a 6 second lead-in leaves room for a retake. Second, extract the frame the lead-in must land on. Third, write the prompt as a description of the motion that ends in that frame, not of the frame itself. Fourth, render the join on the timeline and watch the seam at full speed rather than frame by frame, because a seam that looks clean when paused can still pop in motion.
Sources
Related posts
More in Use cases
- Graduation announcement card art at 4:5, name and year added in code
Generate graduation card art on Sume at 4:5 with a clear middle, then add the graduate's name and year in code so the details are exact and easy to change.
- Halloween safety video with an AI avatar for parents and schools
A 21-second three-scene avatar video built from the FDA Halloween tips: costumes, candy and contact lenses. The request body, the cost by tier, and the label.
- Homebrew beer label art with an API: the name and ABV added in code
Generate square label art on Sume with no text, then draw the beer name, style and ABV in code so every batch gets exact numbers on the same artwork.
- How to split a training lesson into avatar clips under 60 seconds
Sume refuses avatar videos over 60 seconds. Split a 6-minute lesson by words, render each part, and stitch with Timeline 1.0. Word counts and cost by tier.
Written by Sume