Omni 1.1 Flash extends to 40 s; build 40 s on Sume from 4 clips
Google lists extension to 40 s for Omni 1.1 Flash. Sume's gemini-omni-flash-1.1 does 3-10 s per job, so 40 s is four jobs plus a Timeline join.

Google lists scene extension for Gemini Omni 1.1 Flash that reads up to 10 seconds of prior context and reaches 40 seconds in 10-second steps. On Sume, gemini-omni-flash-1.1 accepts 3 to 10 seconds per job, so 40 seconds is four 10-second jobs joined in Timeline, with a still from each clip feeding the next.
That is not the same as Google's extension, which carries prior video context. Sume's chain carries one still.
What Google lists
The Google post (read 2026-10-05) lists scene extension that analyzes up to 10 seconds of prior context, up to 40 seconds total in 10-second steps, first and last frame control, a 360p draft mode, 720p standard and 1080p or 4K premium, and video references up to 3 seconds. It is available in AI Studio, the Gemini API, Google Flow and the Gemini app.
What Sume has for the same id
From the Video Router docs: gemini-omni-flash-1.1 is 3 to 10 seconds at 360p, 720p, 1080p or 4K, in 16:9 or 9:16, with native synced audio always on. Image-to-video takes image_url and an optional end_image_url. Reference-to-video takes up to 10 reference images and up to 3 reference videos of at most 3 seconds each. Edit takes a video_url.
There is no extension capability in that table. The closest tools are the first and last frame fields and the chain.
| Step | Call | Output |
|---|---|---|
| Clip 1 | image_url = opening still, 10 s | 0-10 s |
| Extract | video frames at 9.9 | still for clip 2 |
| Clip 2 | image_url = that still, 10 s | 10-20 s |
| Extract, clip 3, extract, clip 4 | same pattern | 20-40 s |
| Join | Timeline render, 4 slots, 40 s | one MP4 |
Cost shape
Four clips mean four jobs, each billed at the provider list times 1.25 per output second as a function of the resolution. Three frame extractions are billed by their Modal compute. The join is ceil(40/60) = 1 minute, so $0.10. Draft at 360p first to check the story, then regenerate at your final resolution, since the docs show 360p as a supported value.
Make the join invisible
A still-to-clip chain restarts the model every 10 seconds, so the seams are where drift shows. Extract the still a little before the end (for example at 9.9 seconds, since at must be below the duration), keep the prompt style the same, and use a short fade of 0.25 seconds in Timeline 1.0 at the joins. If the model's own audio matters, listen to the seams: audio is always on for this id, and the Timeline spine supplies the audio at assemble time.
If you need one continuous take longer than 10 seconds, the 30-second ids are the better fit: seedance-2.5 and wan-3.0 give 30 seconds in one job.
A first clip request
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-chain-001" \
-d '{"model":"gemini-omni-flash-1.1",
"prompt":"A cyclist rides through a rainy city at dawn",
"image_url":"https://example.com/opening.png",
"duration":10,"resolution":"720p","aspect_ratio":"16:9","mode":"async"}'When a 30-second id is simpler
A 40-second video from 30-second ids is two jobs (30 plus 10) instead of four plus three extractions. If you do not need Omni's 4K option or its edit mode, a 30 second job from wan-3.0 or seedance-2.5 plus a 10 second Omni or other clip joined in Timeline has fewer seams. That shape is the one in the 30 plus 10 second plan.
Draft the story first
Because Omni Flash 1.1 offers 360p, run the four-clip chain at 360p to check pacing and continuity. A bad seam shows at any resolution. Then regenerate only the clips that need to change at 720p or 1080p, and keep the approved ones. Each regenerated clip needs a new Idempotency-Key, or you will get the old job back.
Before you build
Before you build, read the linked Sume docs page for the exact request fields, limits and prices, because those pages are the source of truth and can change. Run one short, cheap test with your own material first, check the output in a player and in your editor, and only then scale to the full shot list. Keep every job id and file you approve, so a later change never forces you to regenerate work that was already signed off. Note that this post describes Sume's catalog and tools; Sume does not run Luma Ray 3.2, and nothing here claims HDR or EXR output.
Sources
Related posts
More in Models
- Omni Flash vs Omni 1.1 Flash: what changed and what Sume exposes
Original Omni Flash was 3-10 s at 720p; 1.1 adds 360p draft, 1080p, 4K, scene extension and frame control. What Sume's id exposes, read 2026-10-05.
- Omni reference limits: 10 images and 3 clips versus Veo's 3 images
On Sume, Gemini Omni takes up to 10 reference images and 3 clips of 3 seconds each; Veo 3.1 takes three images and Lite none. Build the request in Python.
- GPT Image 2.5 can take 2 minutes: submit async, not sync, on Sume
OpenAI says complex GPT Image 2.5 prompts can run up to 2 minutes. Sume's image route blocks only 30 seconds, so send mode async and poll the job.
- GPT Image 2.5 edits through Sume: mask_url and 16 references
Sume's GPT Image 2.5 takes up to 16 references, an optional mask_url and background auto, transparent or opaque. OpenAI: transparent needs png or webp.
Written by Sume