Gemini Omni prompt template: five parts in a Python builder
Google's Omni guide names five prompt parts: framing, style, lighting, location, action. Build them in Python and send the prompt to Sume's /v1/videos.

A Gemini Omni prompt works best when it names five things: shot framing and motion, style, lighting, location, and action. That is the list in Google DeepMind's own prompt guide, and you can enforce it in code by making each part a required argument of one small function before you call Sume's POST /v1/videos with gemini-omni-flash-1.1.
The five parts, straight from the guide
The guide says that adding detail to these five components gives you more control over the output. It also says Omni needs less hand-holding than older models, because it uses reasoning and world knowledge to fill in realistic details. So the five parts are a floor, not a script to pad out.
| Part | Guide wording | Example value |
|---|---|---|
| Framing and motion | Shot framing and motion | Low angle, slow push in |
| Style | Style | Grainy 16mm documentary |
| Lighting | Lighting | Late golden hour, long shadows |
| Location | Location | A fish market on a pier |
| Action | Action | A vendor flips a silver fish onto ice |
Two phrases worth adding
Google's API guide adds two habits. For a single unbroken scene, write "In a single continuous shot" or "No scene cuts". For audio, say what you want out loud, such as calm background music, because Omni produces native audio. Sume's catalog row for gemini-omni-flash-1.1 also says audio is native and always on, so you describe it in the prompt and do not toggle it.
A builder that refuses empty parts
The function below fails fast when a part is missing, then submits with a stable Idempotency-Key so a retry cannot pay twice. The model accepts 3 to 10 seconds, 360p to 4K, and 16:9 or 9:16, per the Sume Video Router docs. Defaults here are 8 seconds at 720p.
import os, time, hashlib, json, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(body):
key = "omni-" + hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()[:24]
r = requests.post("https://api.sume.com/v1/videos", json=body,
headers={**H, "Idempotency-Key": key}, timeout=60)
r.raise_for_status()
url = r.json()["polling_url"]
while True:
s = requests.get(url, headers=H, timeout=30).json()
if s["status"] in ("completed", "failed", "cancelled"):
return s
time.sleep(10)
def build(framing, style, lighting, location, action):
parts = [framing, style, lighting, location, action]
if not all(p.strip() for p in parts):
raise ValueError("all five parts are required")
return ". ".join(parts) + ". In a single continuous shot."
s = run({"model": "gemini-omni-flash-1.1", "duration": 8, "resolution": "720p",
"aspect_ratio": "16:9",
"prompt": build("Low angle, slow push in", "Grainy 16mm documentary",
"Late golden hour", "A fish market on a pier",
"A vendor flips a silver fish onto ice")})
print(s["status"], s.get("usage"), s.get("unsigned_urls"))
What it costs
Sume bills the provider list price times 1.25. The Omni list rate at 720p is $0.10 per second, so 8 seconds is $0.80 list and $1.00 billed. The poll response carries usage.cost, which is the amount Sume billed, so log it next to the prompt that produced it. Sume reserves the estimate at submit and reconciles when the job completes.
A short review checklist
Google's guide also says Omni can reference any kind of media, including images, video and audio, so a weak prompt is often better fixed with a reference image than with more adjectives. On Sume you add those through input_references on /v1/videos.
- Before you spend money on a render, read the prompt back once against this list.
- Does it name a camera move, such as push in or dolly zoom, from the vocabulary in Google's guide?
- Is there exactly one action, so the clip has a single beat inside 8 to 10 seconds?
- Is the audio described, for example room tone or calm music, instead of left to chance?
- Did you keep the first draft at 720p and save 1080p or 4K for the approved prompt?
Iterate on the clip, not the whole prompt
The guide tells you to edit iteratively and says Omni keeps what works across amends. On Sume the edit path is the Video Router's video_url field, which is covered in the Omni edit walkthrough. Keep your five-part function for new shots and use short "keep everything else the same" instructions for edits.
Sources
Related posts
More in Developers
- generate-video 409s: preview_not_ready vs preview_image_not_ready
generate-video on an avatar preview can return two 409s: the job is unfinished, or a public preview image is missing. How to tell them apart.
- generation_spend_cap_usd: null means $500, 0 is a 400, per ad variant
What the per-run spend cap does on Sume Format runs for a number, null, 0 and a value over 500, and why every ad variant should set its own.
- Alert when a new TTS model id lands in the Sume router catalog
September and October brought new voice models. A 20-line script diffs GET /v1/tts-router/models against yesterday's ids and tells you when a row is added.
- GitHub Actions: submit a 30 s video, jobs watch, upload the artifact
A workflow that posts to Sume.s /v1/videos, runs sume jobs watch with a timeout, downloads the clip and uploads it as a build artifact.
Written by Sume