Veo 3.1 prompts cap at 1,024 tokens; Sume's Omni at 20,000 characters
Google caps a Veo 3.1 text prompt at 1,024 tokens. On Sume, gemini-omni-flash-1.1 documents a 20,000-character cap. Tokens and characters differ.

Google's Veo 3.1 page, read 2026-10-01, gives a text input limit of 1,024 tokens. Sume does not list Veo. Its Google video row is gemini-omni-flash-1.1, whose catalog constraint is "prompt ≤20000 characters".
The two limits use different units, so they do not convert exactly. A token is a piece of text that is usually shorter than a word, so count Veo prompts against the token limit and Omni prompts against the character limit.
What fits in a prompt that short?
A one-shot video prompt rarely needs more than subject, action, camera, lighting and sound. If you are hitting a cap, move fixed detail out of the text: send the character or product as a reference image instead of describing it. On Sume, gemini-omni-flash-1.1 takes up to 10 reference images and up to 3 reference videos of 3 seconds each, and addresses them in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, in list order.
What does the call look like?
Reference tokens are 0-based and the reference media is sent before the prompt. One reference image with no first or end frame makes the request reference-to-video. See the video docs for the field list.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-ref-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "<IMAGE_REF_0> walks through a night market, handheld, neon reflections",
"reference_image_urls": ["https://example.com/character.png"],
"resolution": "720p",
"aspect_ratio": "16:9",
"duration": 6
}'How do I check the limit before I submit?
Count characters on your side against the 20,000 cap for Omni. Other Sume models list their own constraints on the model catalog; read the one for the model you pick instead of assuming the Omni number carries over.
How do the two prompt limits line up?
Limits from the sections above.
| Item | Veo 3.1 | Sume gemini-omni-flash-1.1 |
|---|---|---|
| Prompt limit | 1,024 tokens | 20,000 characters |
| Unit | Tokens | Characters |
| Reference images | Per Google's page | Up to 10 |
| Reference videos | Per Google's page | Up to 3, 3 seconds each |
Sources
Related posts
More in Models
- Vidu Q2 Pro Fast image-to-video 1080p vs Sume first frame
QwenCloud lists vidu/viduq2-pro-fast_img2video at 720P and 1080P. Sume has no Vidu id; send image_url as the first frame to a 1080p model.
- Vidu Q3 ad reference-to-video vs Sume reference images
QwenCloud lists a Vidu Q3 ad reference-to-video model. Vidu is not in Sume's video ids; reference images work on models that list them.
- Vidu Q3 drama character consistency vs Sume reference images
QwenCloud lists Vidu Q3 drama for character consistency. Vidu is not in Sume's ids; on Sume, references guide a clip, and frame_images win if both are sent.
- wan2.2-animate-move vs animate-mix, and Sume's motion transfer
animate-move drives a still character with a reference video; animate-mix swaps a character into a video. Sume's Genjutsu row covers the move-style case.
Written by Sume