AI video signs and plates garbled? Write the text in the Omni prompt

Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.

4 min readSume
All posts

If the signs, storefronts or license plates in your AI video come out as nonsense, write the exact text into the prompt. Google's Gemini Omni guide says Omni can render requested text so it is correct and readable, and that if text will appear naturally in the scene, even in background elements, it helps to define what it should say. The same prompt works through Sume's video router, so you do not need a separate step for the text.

Text on screen is where most clips fail a review: a shop sign with invented letters is an instant tell, and a plate or price that contradicts your brief can be a real problem in an ad.

What does Google's own example look like?

The guide gives two prompts. The first sets text that appears one word at a time, and the second plants text on objects in the scene. Quoting the second pattern: a street sign that says "This is an AI generation by Omni", a storefront that says "All you need AI", and a car with the number plate "OMNI1.1".

Text prompt patterns from Google's Gemini Omni guide, read 2026-10-03
Where the text appearsHow the guide words it
On screen, timedOne word on screen at a time, each word for 1s, no dialogue
Street signThere is a street sign that says: "..."
Storefrontthere is a storefront that says: "..."
License platea car with the number plate: "..."
Background elementsDefine what the text should say even when it is incidental

How should I write the text so it survives?

Put each string in quotes and attach it to one object. Keep strings short: a storefront word or two, a plate of a few characters, a label of a few words. Name every piece of text you can see in the frame, so that no sign is left for the model to improvise.

Spell unusual brand names the way you want them to appear and avoid text you do not need. A scene with five signs gives you five chances to be wrong.

How do I send the prompt through Sume?

Use gemini-omni-flash-1.1 on Video Router. It takes 3 to 10 seconds at 360p to 4K in 16:9 or 9:16. Audio is always on, so add "No dialogue" or a music cue if the soundtrack matters. Billing is the provider list rate times 1.25, listed at $0.10 a second at 720p.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-signs-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "A rainy night street, one continuous shot. A shop sign says: \"OPEN LATE\". The storefront says: \"NOODLE BAR\". A parked car has the plate: \"SUME 26\".",
    "resolution": "720p",
    "duration": 6,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

How do I verify the text before I publish?

Do not trust a model's text on one viewing. Pull stills at the moments each sign is in frame with Sume's video frames, which extracts stills at times you name, and read them at full size. Sign text is small, so check the highest resolution you will actually publish: a cheap 360p draft can hide mistakes that show at 1080p.

If one string keeps failing, you have two choices. Regenerate with a shorter or more prominent string, or take the text off the scene and put it on in post. Sume's video captions burns overlay copy onto a video, and its doc says start plus end burns authored overlay copy without speech recognition. That gives you exact spelling for a label or price, at the cost of the text no longer sitting on a physical sign.

Use the in-scene route for atmosphere and the overlay route for anything legal, priced or branded, where one wrong character is not acceptable.

Sources

Related posts

More in Models

All Models posts

Written by Sume