Omni Flash 3-second video reference: trim a clip's last 3 s first
Gemini Omni Flash 1.1 on Sume takes up to 3 video references, each at most 3 s. Trim the tail of a finished clip for $0.02, then cite it as VIDEO_REF_0.

To carry a finished clip's motion into the next shot on Sume's Omni Flash 1.1, trim its last 3 seconds with video trim ($0.02 per job) and send the result as a reference video. The Video Router docs allow up to 3 reference videos, each at most 3 seconds, and you cite them in the prompt as <VIDEO_REF_0>, <VIDEO_REF_1> and so on.
This is reference-to-video, not an extension, so the model is guided by the clip and does not continue it frame by frame. It is one of the closest tools on Sume to the multi-reference idea in Kling 4.0, which lists up to 15 references including up to 5 videos with 30 seconds combined, per the fal explainer (read 2026-10-05).
The two docs pages
The Video Router docs list the reference_to_video mode of gemini-omni-flash-1.1: reference_image_urls up to 10 and reference_video_urls up to 3, each no longer than 3 seconds. The model accepts 3 to 10 seconds, 360p to 4K, 16:9 or 9:16, with native audio always on.
The video trim docs describe the cut: start in seconds, and exactly one of end or duration between 0.2 and 900. The default precision is exact, a frame-accurate re-encode, and the output is a new file in your workspace. A job costs $0.02. Because the trim output is a new file, the original clip is untouched and you can cut a second tail from it later.
Cost of one continuation shot
| Shot | Omni price | Trim | Total |
|---|---|---|---|
| 6 s at 360p | $0.23 | $0.02 | $0.25 |
| 6 s at 720p | $0.75 | $0.02 | $0.77 |
| 6 s at 1080p | $1.13 | $0.02 | $1.15 |
| 6 s at 4K | $2.25 | $0.02 | $2.27 |
What it costs
The price comes from the Omni per-second rate: $0.0375 at 360p, $0.125 at 720p, $0.1875 at 1080p and $0.375 at 4K, so a 6 second clip is $0.23, $0.75, $1.13 and $2.25 at the four tiers, a spread of about ten to one between the smallest and largest. Prices cover the Omni clip only; the trim is a separate $0.02. These are estimates from the pricing package; the final amount is usage.cost on the job.
The cheap way to test the idea is a 360p run. At $0.23 for 6 seconds plus the $0.02 trim, you can check whether the reference changes the result before you pay the 720p price of $0.75, and a handful of 360p tests costs about as much as one 1080p shot.
A reference clip that is the end of the previous shot keeps colour, subject and camera speed consistent. For a hard match on the first frame, use the first-frame route instead: image_url takes a still. You can pull the last frame of the clip with Sume's video frames endpoint.
Trim window and request
This computes the trim window for a clip of known length and builds the two request bodies. It prints them and sends nothing.
import json
clip_seconds = 8
tail = 3
trim = {
"video_url": "https://media.sume.com/artifacts/artf_demo/shot1.mp4",
"start": max(0, clip_seconds - tail),
"duration": min(tail, clip_seconds),
}
generate = {
"model": "gemini-omni-flash-1.1",
"prompt": "Continue the camera move from <VIDEO_REF_0>, same street, later light.",
"reference_video_urls": ["https://media.sume.com/artifacts/artf_demo/tail.mp4"],
"resolution": "720p",
"duration": 6,
"aspect_ratio": "16:9",
}
print(json.dumps(trim))
print(json.dumps(generate, indent=2))Run it in order
Three limits shape the plan. First, the 3 second cap is per reference, so one 3 second tail from the previous shot and up to two more references, such as the opening of an earlier shot, fit in the three video slots. Second, you can add up to 10 image references alongside, for the character or product, and you cite each one as <IMAGE_REF_0> and so on in the same prompt. Third, Omni takes 3 to 10 seconds, so a longer scene is several jobs.
There is no audio reference for this model, and native audio is always on. If you need a fixed voice-over, generate the picture first and lay the audio in a Timeline render afterwards. Keep each shot's Idempotency-Key stable across retries so that a timeout does not create a second paid job.
Run the trim with an Idempotency-Key, wait for result_ready, and take video_url from GET /v1/jobs/{id}/result. Then send the generate body to POST /v1/video-router/generate. The legacy route is the one that documents the flat reference_video_urls field for this model. The video generation docs cover the newer input_references shape.
Sources
Related posts
More in Developers
- Two reference images, one Omni prompt: <IMAGE_REF_0> and <IMAGE_REF_1>
On Sume, Gemini Omni Flash 1.1 takes up to 10 reference images and names them <IMAGE_REF_0>, <IMAGE_REF_1> in list order. Here is a two-reference request.
- A one-file Python CLI that replaces a Sora script, on Sume
One argparse script, standard library only: prompt, model, seconds, resolution and aspect flags in, mp4 out, with an idempotency key derived from the arguments.
- One handler for both: Sume run webhook payload equals the GET receipt
A Sume run webhook's payload is byte-identical to GET /v1/format-runs/{run_id}. Write one function for polling and webhook, and branch on outcome.
- One Sume probe, eight platform limits: a Python preflight
Check one video inspect probe against Truth Social, Odysee, ArtStation, Linktree, Spotlight, Reels, TikTok API and Shorts limits in under 30 lines of Python.
Written by Sume