Gemini Omni Flash on-screen text: the four things to specify
Google's Omni guide says to give text type, placement, animation and exposure. Prompt template, a Sume request, and why to check spelling before publishing.

For on-screen text in Gemini Omni Flash, Google's prompt guide asks you to specify the type of text, its placement, how it animates and how it is exposed. Put those four in the prompt along with the exact words in quotes, then check the spelling on the output, because Google does not promise that rendered text is always correct.
This post uses Google DeepMind's Omni prompt guide and the Gemini API Omni documentation, both read on 2026-10-03. The Sume request uses the catalog id gemini-omni-flash-1.1.
The four fields
Treat the fields as a checklist for every text overlay. Missing any one leaves it to the model.
| Field | What to write | Example wording |
|---|---|---|
| Type | Kind of text and style | bold white title text, lower-third caption |
| Placement | Where in frame | centered, top left, bottom third |
| Animation | How it appears | fades in, slides up, types on |
| Exposure | How long and when | visible from 1s to 4s |
Keep the words short
Quote the text exactly and keep it to a few words. Long lines give the model more chances to misspell. If the clip needs a full sentence, a better path is to render a clean clip and burn captions afterward with Sume's caption job, which draws text from a transcript rather than from a video model. That route is covered in a related post on readable text.
Note that the caption job needs spoken words: a silent clip fails with caption_no_speech. Omni's native audio can provide them if you ask for speech.
Request
The text spec goes in the same prompt as the scene.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-text-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "In a single continuous shot, a coffee cup on a wooden table, steam rising. Bold white title text reading \"OPEN LATE\" centered in the top third, fades in at 1 second and stays visible until 4 seconds.",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "9:16",
"mode": "async"
}'Proof the output
Sample stills at the timestamps where the text should be visible and read them yourself. Look for swapped letters, doubled characters and text that appears outside the window you asked for. Fix problems by shortening the text or moving it to a plainer area of the frame, then rerun. The Video Router docs list what the endpoint accepts.
Sources
Related posts
More in Models
- Gemini Omni Flash adds cuts: how to prompt a single shot
Omni Flash tries a few shots by default. Google's wording for one unbroken take, and the same request on Sume's Video Router with a check for hidden cuts.
- Choose the music in a Gemini Omni clip by prompting the audio
Gemini Omni makes its own soundtrack. Google's guide shows how to steer it with a music style, a radio effect or a timed chorus; here is a cheap way to test.
- Gemini Omni rapid-fire video: a new labelled item every second
Prompt Gemini Omni for a rapid-fire clip that shows a different item every second with a text label, then send it through Sume's video router in 9:16.
- AI video signs and plates garbled? Write the text in the Omni prompt
Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.
Written by Sume