Gemini Omni Flash adds cuts: how to prompt a single shot
Omni Flash tries a few shots by default. Google's wording for one unbroken take, and the same request on Sume's Video Router with a check for hidden cuts.

Gemini Omni Flash tries to build a short narrative from a prompt, so a clip often contains a few different shots. To get one unbroken take, say so in the prompt: Google's guide gives "In a single unbroken scene", "In a single continuous shot" and "No scene cuts" as phrases that work. Put the phrase first, then describe the camera and the action once.
These details come from Google DeepMind's Omni prompt guide and the Gemini API Omni documentation, both read on 2026-10-03. The Sume request at the end uses the catalog id gemini-omni-flash-1.1.
Phrases Google's guide gives
The guide separates three jobs: forcing one scene, naming a continuous shot style, and fixing the camera. The first two stop automatic cutting; the third stops camera drift. Use one from each group, not all of them at once, because a long list of constraints tends to read as noise.
| Goal | Wording the guide lists |
|---|---|
| Force one scene | "In a single unbroken scene", "In a single continuous shot", "No scene cuts" |
| Continuous take style | "one continuous shot" or "oner" |
| Camera that does not move | "static", "locked off", "fixed" |
| Moving camera | "push in", "punch in", "dolly zoom" |
Where to put the phrase
Start the prompt with the shot instruction and keep the action to one clear beat per second or two. A prompt like "In a single continuous shot, locked off, a barista pours milk into a cup and slides it across the counter" gives the model one job. Adding "then cut to" or a second location invites exactly the cut you tried to avoid.
If you actually want cuts at known moments, the other direction is also documented in existing Sume posts such as the one on timecode prompts for scene cuts. Choose one approach per clip.
Check the result for hidden cuts
A short clip can hide a cut in a fast move. Sume's Video inspect surface probes a Sume-hosted clip and samples stills from it, so you can request frames across the clip and look for a change of setting. It reads one media.sume.com clip, never re-encodes it, and is billed by its compute, so it costs far less than another render. Sample at least one still per second of a 10 second clip.
The request
Omni on Sume accepts 3 to 10 seconds at 360p, 720p, 1080p or 4K, in 16:9 or 9:16, with synced audio always on. The Idempotency-Key header makes a retry safe.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-single-shot-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "In a single continuous shot, locked off, a barista pours milk into a cup and slides it across a wooden counter. Morning window light.",
"resolution": "720p",
"duration": 6,
"aspect_ratio": "16:9",
"mode": "async"
}'Sources
Related posts
More in Models
- Choose the music in a Gemini Omni clip by prompting the audio
Gemini Omni makes its own soundtrack. Google's guide shows how to steer it with a music style, a radio effect or a timed chorus; here is a cheap way to test.
- Gemini Omni rapid-fire video: a new labelled item every second
Prompt Gemini Omni for a rapid-fire clip that shows a different item every second with a text label, then send it through Sume's video router in 9:16.
- AI video signs and plates garbled? Write the text in the Omni prompt
Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.
- German and Italian text to speech API: de and it on Sume TTS
German (de) and Italian (it) are in both Cartesia Sonic 3.6 and Sume's voice library. Send language de or it, reuse one voice, and watch the 409 language check.
Written by Sume