Gemini Omni rapid-fire video: a new labelled item every second
Prompt Gemini Omni for a rapid-fire clip that shows a different item every second with a text label, then send it through Sume's video router in 9:16.

To get a rapid-fire clip from Gemini Omni, say the rhythm in the prompt: one item per interval, a music cue, and a request to label each item. Google's Omni guide gives this as a meta prompt: "Make a rapid fire video that shows a different rare [thing] every 1s, upbeat music, include text to label the thing." On Sume the clip is 3 to 10 seconds, so a one-second rhythm gives you at most about ten items.
This is the format behind many list-style Shorts, such as ten rare fruits or ten pieces of desk gear. The model does the cutting and the on-screen labels, so you need one request and no editor.
What prompt wording does Google document?
The guide has a "Timing events" section. It says there is no precise syntax needed and natural language works, and it lists these examples.
- "Every 2s cut to a new frame."
- "In a rapid fire sequence, every half a second (12 frames at 24fps) change the scene to a new location."
- "After 3 seconds, a woman enters the scene."
- A timecode form: [0-3s] A person is walking, [3-6s] They stop and turn around.
How many items fit in one clip?
Divide the clip length by the interval. The guide's own half-second example is written as 12 frames at 24 fps. Sume's catalog allows 3 to 10 seconds per gemini-omni-flash-1.1 clip, so the table shows the arithmetic. The item counts are simple division, not a guarantee that the model lands every cut exactly.
| Clip length | Every 2 s | Every 1 s | Every 0.5 s |
|---|---|---|---|
| 3 s | 1 to 2 | 3 | 6 |
| 6 s | 3 | 6 | 12 |
| 10 s | 5 | 10 | 20 |
How do I send it for a vertical Short?
Sume's Video Router takes gemini-omni-flash-1.1 with resolution, duration and aspect_ratio set to 16:9 or 9:16. Native synced audio is always on, so the "upbeat music" in the prompt is the only way to steer the soundtrack; there is no reference audio field. Billing is the provider list price times 1.25 per output second, and the list rate at 720p is $0.10 a second, so a 10-second 720p clip lists at $1.00 before the multiplier.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-rapid-fire-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Make a rapid fire video that shows a different rare fruit every 1s, upbeat music, include text to label the fruit.",
"resolution": "720p",
"duration": 10,
"aspect_ratio": "9:16",
"mode": "async"
}'What can go wrong, and how do I check it?
Rapid cuts are where a model is most likely to drift: a label that disagrees with the picture, or two items sharing one second. Google's guide says the model renders requested text in a way that is correct and readable, and that spelling out what text should say helps. So name the items in the prompt instead of writing "a different fruit" and hoping.
Before you spend on 1080p, review a cheap draft. Sume's video frames extracts stills at times you name, so ask for one frame at the middle of each second and compare the label to the picture. A 360p draft is listed at $0.03 a second on the same API doc, which makes a first pass cheap.
If the cuts land but the music is wrong, change the audio wording and regenerate: Google's guide suggests describing the audio you want, for example a high-energy techno beat. If you want a different bed, Sume's Timeline accepts an optional soundtrack with gain_db, loop and duck_db fields, so you can lay music under the finished clip. Check what the render did with the clip's own sound by running video inspect on the result.
Sources
Related posts
More in Models
- AI video signs and plates garbled? Write the text in the Omni prompt
Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.
- German and Italian text to speech API: de and it on Sume TTS
German (de) and Italian (it) are in both Cartesia Sonic 3.6 and Sume's voice library. Send language de or it, reuse one voice, and watch the 409 language check.
- Google Pics API? Edit one object with Nano Banana on Sume
Google's Pics announcement describes an app, not an API. The closest call on Sume: Nano Banana reference edits, or a GPT Image 2.5 mask for one region.
- Google video model dates: Omni GA, Veo shutdowns, one table
One dated table of Google's video model lifecycle read from its own pages on Oct 3, 2026: Omni 1.1 Flash GA, Veo 3.1 preview shutdowns, and what to call.
Written by Sume