Veo 3.1 Lite 1,024-token prompt limit vs Omni's 20,000 characters
Google's Veo 3.1 Lite page caps text input at 1,024 tokens. Sume's gemini-omni-flash-1.1 row takes 20,000 characters. What that means for long shot lists.

The Gemini API's Veo 3.1 Lite model page lists a text input limit of 1,024 tokens, with text and image input and video with audio output (read 2026-10-04). That is room for a tight single-shot prompt, not a full script with timing marks.
Sume does not list Veo 3.1 Lite, but it does list a Google video model: gemini-omni-flash-1.1. Its catalog row sets a prompt limit of 20,000 characters, 3 to 10 second clips, ratios 16:9 and 9:16, and resolutions up to 4K. Tokens and characters are different units, so the two numbers are not a strict ratio, but the Omni row leaves far more space for detail.
When prompt length matters
If your prompt is short, the limit never bites and price decides: the Gemini pricing page, read 2026-10-04, lists Veo 3.1 Lite at $0.05 per second at 720p and $0.08 at 1080p, while Sume bills Omni 1080p at $0.1875 per second.
The Gemini API pricing page is the source for the Veo numbers.
- Shot-by-shot descriptions with camera moves for each second of a 10-second clip.
- Brand rules, character descriptions and negative instructions pasted in with the shot.
- Dialogue lines with timing, when audio is generated with the clip.
Send a long prompt
The video generation docs show the body. This request puts a multi-line prompt on Omni.
import os
import requests
prompt = "\n".join([
"0-3s: wide shot of a night market, handheld, warm neon.",
"3-6s: push in on a vendor flipping skewers, steam rising.",
"6-10s: close on a customer smiling, shallow focus.",
])
body = {
"model": "gemini-omni-flash-1.1",
"prompt": prompt,
"duration": 10,
"resolution": "1080p",
"aspect_ratio": "9:16",
}
r = requests.post(
"https://api.sume.com/v1/videos",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "omni-long-001",
},
json=body,
)
print(r.status_code)Sources
Related posts
More in Models
- Veo 3.1 Lite takes text and image only; Sume models take more inputs
Google lists Veo 3.1 Lite inputs as text and image. On Sume, Seedance 2.0 and Wan 3.0 also accept video and audio references; Kling 3 does not.
- Video aspect ratios by Sume row: Kling 3, Gemini, Wan 3.0, MiniMax H3
Which aspect_ratio values each Sume video row takes: kling-3 has three, Gemini two, Wan and MiniMax add adaptive, and MiniMax adds 21:9. From the catalog.
- MAI-Voice-2.1-Flash tops out at 45 seconds: what for longer?
MAI-Voice-2.1-Flash is reported at up to 45 s of audio and 150 ms latency. Longer lines need the standard model or a TTS taking 20,000 characters.
- Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.
Written by Sume