AI video signs and plates garbled? Write the text in the Omni prompt
Gemini Omni renders text well when you say what it reads. Google's guide covers signs, storefronts and plates; here is the prompt pattern and a Sume request.

If the signs, storefronts or license plates in your AI video come out as nonsense, write the exact text into the prompt. Google's Gemini Omni guide says Omni can render requested text so it is correct and readable, and that if text will appear naturally in the scene, even in background elements, it helps to define what it should say. The same prompt works through Sume's video router, so you do not need a separate step for the text.
Text on screen is where most clips fail a review: a shop sign with invented letters is an instant tell, and a plate or price that contradicts your brief can be a real problem in an ad.
What does Google's own example look like?
The guide gives two prompts. The first sets text that appears one word at a time, and the second plants text on objects in the scene. Quoting the second pattern: a street sign that says "This is an AI generation by Omni", a storefront that says "All you need AI", and a car with the number plate "OMNI1.1".
| Where the text appears | How the guide words it |
|---|---|
| On screen, timed | One word on screen at a time, each word for 1s, no dialogue |
| Street sign | There is a street sign that says: "..." |
| Storefront | there is a storefront that says: "..." |
| License plate | a car with the number plate: "..." |
| Background elements | Define what the text should say even when it is incidental |
How should I write the text so it survives?
Put each string in quotes and attach it to one object. Keep strings short: a storefront word or two, a plate of a few characters, a label of a few words. Name every piece of text you can see in the frame, so that no sign is left for the model to improvise.
Spell unusual brand names the way you want them to appear and avoid text you do not need. A scene with five signs gives you five chances to be wrong.
How do I send the prompt through Sume?
Use gemini-omni-flash-1.1 on Video Router. It takes 3 to 10 seconds at 360p to 4K in 16:9 or 9:16. Audio is always on, so add "No dialogue" or a music cue if the soundtrack matters. Billing is the provider list rate times 1.25, listed at $0.10 a second at 720p.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-signs-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A rainy night street, one continuous shot. A shop sign says: \"OPEN LATE\". The storefront says: \"NOODLE BAR\". A parked car has the plate: \"SUME 26\".",
"resolution": "720p",
"duration": 6,
"aspect_ratio": "9:16",
"mode": "async"
}'How do I verify the text before I publish?
Do not trust a model's text on one viewing. Pull stills at the moments each sign is in frame with Sume's video frames, which extracts stills at times you name, and read them at full size. Sign text is small, so check the highest resolution you will actually publish: a cheap 360p draft can hide mistakes that show at 1080p.
If one string keeps failing, you have two choices. Regenerate with a shorter or more prominent string, or take the text off the scene and put it on in post. Sume's video captions burns overlay copy onto a video, and its doc says start plus end burns authored overlay copy without speech recognition. That gives you exact spelling for a label or price, at the cost of the text no longer sitting on a physical sign.
Use the in-scene route for atmosphere and the overlay route for anything legal, priced or branded, where one wrong character is not acceptable.
Sources
Related posts
More in Models
- German and Italian text to speech API: de and it on Sume TTS
German (de) and Italian (it) are in both Cartesia Sonic 3.6 and Sume's voice library. Send language de or it, reuse one voice, and watch the 409 language check.
- Google Pics API? Edit one object with Nano Banana on Sume
Google's Pics announcement describes an app, not an API. The closest call on Sume: Nano Banana reference edits, or a GPT Image 2.5 mask for one region.
- Google video model dates: Omni GA, Veo shutdowns, one table
One dated table of Google's video model lifecycle read from its own pages on Oct 3, 2026: Omni 1.1 Flash GA, Veo 3.1 preview shutdowns, and what to call.
- GPT-5.1, o3, GPT-5.4 Nano shutdown dates and the Sume model field
OpenAI lists shutdowns for gpt-5.1, gpt-5.3-codex, gpt-5.4-nano (Apr 1, 2027) and o3 (Dec 11, 2026). A Sume Format run with those ids gets a 400.
Written by Sume