Split a 40-second brief into four Omni prompts, one character block
A Python script that turns one 40-second brief into four 10-second Gemini Omni 1.1 Flash request bodies for Sume, with one shared character block.

To turn a 40-second brief into four Gemini Omni 1.1 Flash jobs on Sume, write one character block, repeat it word for word at the start of every prompt, and give each prompt its own 10-second action. Each job is a text-to-video request of 10 seconds. Google's own extension feature moves in 10-second steps to 40 seconds (Google, read 2026-10-05); on Sume you do the same split yourself, since the catalog model takes 3 to 10 seconds per job (Sume docs: Video Router, read 2026-10-05).
Why a shared block
Four separate prompts are four separate generations. Nothing carries a face or a coat from one to the next unless the text, or a reference, does. A fixed character block, with the same words each time, is the cheapest way to keep a person recognisable. It does not guarantee a match. For a stronger tie, pass a still as reference_image_urls and name it <IMAGE_REF_0> in the prompt.
Write the block as concrete visible facts: age band, hair, clothing colours, one prop. Avoid adjectives that do not show on screen, such as confident or friendly. Two or three clothing colours are enough; each extra detail is another thing the model can change between shots.
The script
It builds the four bodies and prints them. It makes no network call, so it runs as written. Submit each body to POST /v1/videos with its own Idempotency-Key.
import json
CHARACTER = ("A woman in her thirties, short black hair, mustard raincoat, "
"round glasses, carrying a small green backpack.")
STYLE = "Handheld, overcast daylight, muted colours, natural sound."
BEATS = [
"She leaves a train station and checks a paper map.",
"She walks through a street market past fruit stalls.",
"She climbs a long stair beside a canal.",
"She reaches a rooftop and looks out over the city.",
]
def bodies(resolution="360p"):
return [{
"model": "gemini-omni-flash-1.1",
"prompt": f"{CHARACTER} {beat} {STYLE}",
"duration": 10,
"resolution": resolution,
"aspect_ratio": "16:9",
} for beat in BEATS]
if __name__ == "__main__":
for i, b in enumerate(bodies()):
print(f"shot-{i + 1}", json.dumps(b))Using it
Note that the finish pass is a new generation, not an upscale of the draft: a 720p run of the same prompt can give a different take. If you need the same take, use the first frame of the approved draft as image_url, or accept the new take.
Order the beats so each one ends where the next one starts. If beat one ends with her at the market entrance, beat two should begin in the market. That is a cheap form of continuity, and it costs nothing but the sentence.
Idempotency keys should name the shot and the version, such as trailer-shot-2-v1. When you change a beat, bump the version; when a network call fails, resend the same key and the original job is replayed instead of billed twice.
Finally, keep the output of the script in a file next to the clips. The four bodies are your shot list, and they let anyone rebuild the sequence at 720p later without reading chat history.
Budget before you submit. Four 10-second drafts are 40 output seconds at the 360p rate, and four finals are 40 seconds at the 720p rate. Multiply by the live per-second price from the models endpoint, add the timeline's $0.10 for one minute, and compare the total with what you meant to spend. If it is too high, shorten the beats to 6 or 8 seconds before you run anything.
- Run it at
360pfirst. Look at all four, and re-roll only the shot that breaks the character. - Change
resolutionto720pand run the keepers again for the finish. A 360p draft is described by Google as a third of the cost of a 720p one. - Join the four results with Sume's timeline. The timeline plan post shows the unbilled plan call for four 10-second slots.
- Keep the character block in one variable. If you edit it between shots, the shots drift.
Sources
Related posts
More in Developers
- Spread Graph API calls evenly: pace a nightly Reel batch
Meta advises spreading queries evenly to avoid traffic spikes. Space publish calls across the hour, and size the Sume render wave from generation_limits.
- SQLite ledger: a restarted poller never resubmits a 30-second video
Save the Sume job id and Idempotency-Key in SQLite before you poll, so a crash resumes the same $17 render instead of paying twice. Python, stdlib.
- Square YouTube Short at 1080x1080: set Timeline output size
YouTube classifies square or vertical videos up to 3 minutes as Shorts. Set Timeline 1.0 output width and height to 1080 for square; 1080x1920 for vertical.
- Startup plan accepts 48 jobs, not 50: send 50 AI video clips in waves
Sume admits concurrency plus queue jobs per plan: 6, 24, 48 or 120. Fifty 30-second clips fit only on Scale at once. See waves per plan and how to retry.
Written by Sume