Which Omni input continues a scene: end frame, reference clip or edit?
Sume has no extend button. Continue an Omni scene with a last-frame image, a 3 s reference clip, or a video edit; this table says which, with costs per 10 s.

On Sume, continue an Omni scene by sending the last frame of the previous clip as image_url for a continuation, or the last three seconds as reference_video_urls to carry the look; use video_url only when you mean to change the clip you already have. The three inputs do different jobs, and only the first two make new footage.
Google describes scene extension in Omni 1.1 Flash as analyzing up to 10 seconds of prior context and extending in 10-second steps up to 40 seconds (Google blog, read 2026-10-05). The Video Router page says Sume does not expose previous_interaction_id or extend, so a continuation on Sume is a new job whose inputs you choose.
The three inputs side by side
All three go to gemini-omni-flash-1.1, which routes by the shape of the request. Rules come from the Video Router page.
| Input | Field | What it does | Limits |
|---|---|---|---|
| Last frame as the opener | image_url (+ optional end_image_url) | New clip that starts from a still | 3-10 s, 16:9 or 9:16 |
| Reference clip | reference_video_urls | New clip guided by the look of the footage | up to 3 clips, each up to 3 s |
| Edit | video_url | Rewrites the clip you send | No aspect_ratio or duration; source sets the output |
Option 1: last frame as the opener
Extract the final still of the previous clip with the Video frames API, then send it as image_url with a prompt that says what happens next. A still has position and light but not motion, so write the motion into the prompt: "the camera keeps drifting right as she opens the door". This is the most predictable route for a hard continuation, and it works for 3 to 10 seconds.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: last-frame-017" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip1.mp4",
"at": [9.9]
}'Option 2: a reference clip
Trim the last 3 seconds with the Video Trim API, at $0.02 per job, and send the result in reference_video_urls. In the prompt, address it as <VIDEO_REF_0>. Video references carry the look and character but, like any reference, do not fix the first frame of the new clip, so the cut may not be seamless. Google caps video references at up to 3 seconds as well.
Option 3: edit
video_url is for changing what you have: swap an object, change the light. It does not extend. If the clip is too short, edit is the wrong tool, because the output length follows the source.
What a 40-second scene costs this way
Four 10-second clips at 720p are $5.00, at 1080p $7.52, at 4K $15.00, plus about $0.02 per trim. The continuity depends on your prompts and on the chosen input, not on a stored context.
| Resolution | Per 10 s clip | Four clips |
|---|---|---|
| 360p | $0.38 | $1.52 |
| 720p | $1.25 | $5.00 |
| 1080p | $1.88 | $7.52 |
| 4K | $3.75 | $15.00 |
Common mistakes
The first mistake is sending the previous clip as video_url and expecting a longer clip. The edit route returns the same length as the source and rejects aspect_ratio, so you get a changed copy, not a continuation. The second is mixing video_url with image_url or reference fields; the API rejects the combination. The third is sending a reference clip longer than 3 seconds. Trim it first.
A fourth, quieter mistake is changing the aspect ratio between clips. Omni takes 16:9 or 9:16 only, so pick one for the whole scene and send it on every request.
Keep the audio in mind
Native audio is always on for this model, so each new clip has its own generated sound. A joined scene can have audible steps at the cuts. Plan to lay one audio spine under the joined clips on a Timeline, or accept the steps as scene changes.
Pick by the problem
- Need the exact end position to carry on: last frame as opener.
- Need the same character and style, any pose: reference clip.
- Need the same clip with one change: edit.
- Need a long unbroken take: check the length limits of other models in the Video Router before chaining 10-second clips.
Sources
Related posts
More in Developers
- What ends a Sume STT sentence segment: . ! ? and the Japanese marks
Sume ends a segment on . ! ? 。 ! ? … plus optional trailing quotes or closing parentheses; the corner bracket 」 and a fullwidth ) are not on that list.
- Which API key scope does each Sume webhook endpoint need?
The signing secret needs account:read, rotate and test deliveries need account:write, redeliver needs jobs:write or formats:write. Map each call to a key.
- Which Sume API requests count against the read rate limit?
Every GET and HEAD is a read, and so are POST /v1/generation/admission-preview and the MCP endpoint. Reads have their own per-key bucket, 40x the write one.
- Which Sume API routes are missing from the public OpenAPI document?
Asset upload routes, admission-preview, POST /v1/avatars, POST /v1/avatar-videos and the generic model runs path are implemented but hidden from OpenAPI.
Written by Sume