Omni scene extension on Sume: chain clips with a 3-second tail
Google's extension reads 10 seconds of prior context. Sume's Omni row takes video references of 3 seconds. Cut the tail with video-trim and chain clips by hand.

Can you chain Gemini Omni clips on Sume to approximate a long scene? Partly. Google's extension reads up to 10 seconds of prior context, but Sume's gemini-omni-flash-1.1 row accepts reference videos of at most 3 seconds each, up to 3 of them. So you cut the last 3 seconds of clip N with video-trim, pass it as a video reference to clip N+1, and join the full clips afterward.
The numbers on each side
Google's launch post says prior context grew from 1 second to up to 10 seconds, with extension in 10 second increments up to 40 seconds. The Gemini API docs give video references as up to 3 clips of up to 3 seconds.
| Route | Prior context | Max length |
|---|---|---|
| Google scene extension | Up to 10 s | 40 s total |
| Sume reference videos | 3 clips, 3 s each | 3-10 s per clip, chained by you |
The chain
Take a finished clip that lives on media.sume.com. Send POST /v1/video-trim with video_url, a start three seconds before the end, and a duration of 3, plus an Idempotency-Key. The video-trim docs list precision: exact as the default, a frame-accurate re-encode, and a rate of $0.02 per job. Poll the job until its result gives a new video_url.
Then submit the next generation with that URL as a video entry in input_references, and describe in the prompt what happens next. Repeat for each segment.
Where it is weaker than real extension
Three seconds of context carries motion direction and look, but not a 10 second memory of the scene, so characters and props can drift over several links. Keep each prompt specific about who is in frame. Also, Sume does not expose previous_interaction_id for Omni, so there is no server-side conversation state; each clip is a fresh request.
To assemble the pieces, render all the full clips in order on a timeline, not the trimmed tails, which only serve as references. Check the seam between clips at full speed before you accept the chain.
A worked plan for 30 seconds
Three 10 second Omni clips make a 30 second piece. Generate clip one from your prompt and a first frame. Trim its last 3 seconds and use them as the reference for clip two. Repeat for clip three. Total cost is three generations plus two trims at $0.02 each, plus the final timeline render. At the 720p catalog rate, each 10 second clip bills $1.25 on Sume, so the generations come to $3.75 before trims.
Compare that against a single Seedance 2.5 or Wan 3.0 request, which can reach 30 seconds in one pass, and decide whether you need Omni's native audio or look enough to justify the chain.
Sources
Related posts
More in Developers
- One API key for Seedance 2.5, Wan 3.0, Kling 3 and MiniMax H3
Do you need a ByteDance, Alibaba, Kuaishou and MiniMax account to call their video models? On Sume one key and one wallet cover all four. How it works.
- One 6-minute b-roll, twelve Shorts episodes: source_in offsets
Slice one imported b-roll into twelve non-repeating 30-second episode backgrounds with Timeline 1.0 source_in, and plan the whole season unbilled in Python.
- One Sume API key per service: what it isolates and what it does not
Sume request budgets are per key and reads and writes are already separate. A key per service isolates revocation and scope, not generation capacity.
- OpenAI images.generate to Sume /v1/images: field by field map
Move a gpt-image-1 images.generate call to Sume POST /v1/images: which fields carry over, which return 400, and why size becomes image_size. Python mapper.
Written by Sume