Gemini Omni Flash on-screen text: the four things to specify

Google's Omni guide says to give text type, placement, animation and exposure. Prompt template, a Sume request, and why to check spelling before publishing.

5 min readSume
All posts

For on-screen text in Gemini Omni Flash, Google's prompt guide asks you to specify the type of text, its placement, how it animates and how it is exposed. Put those four in the prompt along with the exact words in quotes, then check the spelling on the output, because Google does not promise that rendered text is always correct.

This post uses Google DeepMind's Omni prompt guide and the Gemini API Omni documentation, both read on 2026-10-03. The Sume request uses the catalog id gemini-omni-flash-1.1.

The four fields

Treat the fields as a checklist for every text overlay. Missing any one leaves it to the model.

Text prompt checklist from Google's Omni guide (read 2026-10-03)
FieldWhat to writeExample wording
TypeKind of text and stylebold white title text, lower-third caption
PlacementWhere in framecentered, top left, bottom third
AnimationHow it appearsfades in, slides up, types on
ExposureHow long and whenvisible from 1s to 4s

Keep the words short

Quote the text exactly and keep it to a few words. Long lines give the model more chances to misspell. If the clip needs a full sentence, a better path is to render a clean clip and burn captions afterward with Sume's caption job, which draws text from a transcript rather than from a video model. That route is covered in a related post on readable text.

Note that the caption job needs spoken words: a silent clip fails with caption_no_speech. Omni's native audio can provide them if you ask for speech.

Request

The text spec goes in the same prompt as the scene.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-text-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "In a single continuous shot, a coffee cup on a wooden table, steam rising. Bold white title text reading \"OPEN LATE\" centered in the top third, fades in at 1 second and stays visible until 4 seconds.",
    "resolution": "720p",
    "duration": 5,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

Proof the output

Sample stills at the timestamps where the text should be visible and read them yourself. Look for swapped letters, doubled characters and text that appears outside the window you asked for. Fix problems by shortening the text or moving it to a plainer area of the frame, then rerun. The Video Router docs list what the endpoint accepts.

Sources

Related posts

More in Models

All Models posts

Written by Sume