Gemini Omni prompt template: five parts in a Python builder

Google's Omni guide names five prompt parts: framing, style, lighting, location, action. Build them in Python and send the prompt to Sume's /v1/videos.

4 min readSume
All posts

A Gemini Omni prompt works best when it names five things: shot framing and motion, style, lighting, location, and action. That is the list in Google DeepMind's own prompt guide, and you can enforce it in code by making each part a required argument of one small function before you call Sume's POST /v1/videos with gemini-omni-flash-1.1.

The five parts, straight from the guide

The guide says that adding detail to these five components gives you more control over the output. It also says Omni needs less hand-holding than older models, because it uses reasoning and world knowledge to fill in realistic details. So the five parts are a floor, not a script to pad out.

Prompt parts named in the Gemini Omni prompt guide (read 2026-10-05)
PartGuide wordingExample value
Framing and motionShot framing and motionLow angle, slow push in
StyleStyleGrainy 16mm documentary
LightingLightingLate golden hour, long shadows
LocationLocationA fish market on a pier
ActionActionA vendor flips a silver fish onto ice

Two phrases worth adding

Google's API guide adds two habits. For a single unbroken scene, write "In a single continuous shot" or "No scene cuts". For audio, say what you want out loud, such as calm background music, because Omni produces native audio. Sume's catalog row for gemini-omni-flash-1.1 also says audio is native and always on, so you describe it in the prompt and do not toggle it.

A builder that refuses empty parts

The function below fails fast when a part is missing, then submits with a stable Idempotency-Key so a retry cannot pay twice. The model accepts 3 to 10 seconds, 360p to 4K, and 16:9 or 9:16, per the Sume Video Router docs. Defaults here are 8 seconds at 720p.

import os, time, hashlib, json, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(body):
    key = "omni-" + hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()[:24]
    r = requests.post("https://api.sume.com/v1/videos", json=body,
                      headers={**H, "Idempotency-Key": key}, timeout=60)
    r.raise_for_status()
    url = r.json()["polling_url"]
    while True:
        s = requests.get(url, headers=H, timeout=30).json()
        if s["status"] in ("completed", "failed", "cancelled"):
            return s
        time.sleep(10)

def build(framing, style, lighting, location, action):
    parts = [framing, style, lighting, location, action]
    if not all(p.strip() for p in parts):
        raise ValueError("all five parts are required")
    return ". ".join(parts) + ". In a single continuous shot."

s = run({"model": "gemini-omni-flash-1.1", "duration": 8, "resolution": "720p",
         "aspect_ratio": "16:9",
         "prompt": build("Low angle, slow push in", "Grainy 16mm documentary",
                         "Late golden hour", "A fish market on a pier",
                         "A vendor flips a silver fish onto ice")})
print(s["status"], s.get("usage"), s.get("unsigned_urls"))

What it costs

Sume bills the provider list price times 1.25. The Omni list rate at 720p is $0.10 per second, so 8 seconds is $0.80 list and $1.00 billed. The poll response carries usage.cost, which is the amount Sume billed, so log it next to the prompt that produced it. Sume reserves the estimate at submit and reconciles when the job completes.

A short review checklist

Google's guide also says Omni can reference any kind of media, including images, video and audio, so a weak prompt is often better fixed with a reference image than with more adjectives. On Sume you add those through input_references on /v1/videos.

  • Before you spend money on a render, read the prompt back once against this list.
  • Does it name a camera move, such as push in or dolly zoom, from the vocabulary in Google's guide?
  • Is there exactly one action, so the clip has a single beat inside 8 to 10 seconds?
  • Is the audio described, for example room tone or calm music, instead of left to chance?
  • Did you keep the first draft at 720p and save 1080p or 4K for the approved prompt?

Iterate on the clip, not the whole prompt

The guide tells you to edit iteratively and says Omni keeps what works across amends. On Sume the edit path is the Video Router's video_url field, which is covered in the Omni edit walkthrough. Keep your five-part function for new shots and use short "keep everything else the same" instructions for edits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume