Veo 3.1 Lite takes text and images only: no references or extension
Google's Veo page shows Lite without reference images, video input or extension. If your Lite job needs those, here is what to use instead on Sume.

Veo 3.1 Lite accepts text and an image, with an optional last frame, and nothing else. Google's Veo page lists Lite's input modalities as text-to-video and image-to-video, its Lite model card gives the input as text and image, and it says extension is for Veo 3.1 and Veo 3.1 Fast only. So a Lite job that needs a product reference, a character sheet, or a clip to extend has to move to a different model.
That move is coming anyway. Google's deprecations page lists the Lite preview id with a shutdown date of October 22, 2026 and names gemini-omni-1.1-flash as the replacement.
What can each model take as input?
The table compares the Veo parameters table with what Sume's Video Router doc says about Gemini Omni Flash 1.1. Sume's catalog does not list Veo, so the Omni column is the only Google-family model you can call there.
| Input | Veo 3.1 / 3.1 Fast | Veo 3.1 Lite | Sume gemini-omni-flash-1.1 |
|---|---|---|---|
| Text prompt | Yes | Yes | Yes (prompt) |
| Start image | Yes | Yes | Yes (image_url) |
| Last frame | Yes | Yes | Yes (end_image_url) |
| Reference images | Up to three | Not listed | Up to 10 (reference_image_urls) |
| Reference video | Not listed | Not listed | Up to 3 clips, each 3 s or less |
| Video input to extend | Veo-made clips only | Not available | Not exposed |
| Video input to edit | Listed as video-to-video | Not listed | Yes (video_url) |
What do I use for references on Sume?
Send up to ten reference_image_urls and, if you want motion or likeness, up to three reference_video_urls, each no longer than three seconds. Address them in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>, counting from zero in list order. Google's Veo page caps reference images at three and requires an 8-second duration when you use them, while Sume's Omni row keeps its normal 3 to 10 second range.
One caution from Google's guide that applies on both sides: video references work best with likenesses, any audio in a reference video is ignored, and referencing across multiple videos can degrade results.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-refs-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "The person in <IMAGE_REF_0> holds the bottle in <IMAGE_REF_1> on a sunny balcony.",
"reference_image_urls": [
"https://example.com/person.jpg",
"https://example.com/bottle.jpg"
],
"resolution": "720p",
"duration": 6,
"aspect_ratio": "9:16",
"mode": "async"
}'What about extending a clip?
Neither side has a simple answer. Google's Veo page limits extension to Veo-made clips and to Veo 3.1 or 3.1 Fast. Sume's Omni row does not expose the extend task or previous_interaction_id; its doc says they are not on the provider schema Sume calls. If you need a longer shot, the Video Router catalog has models with longer limits, for example seedance-2.5 at 4 to 30 seconds and wan-3.0 at 2 to 30 seconds, or you can chain clips with Timeline.
For an existing clip you want changed rather than lengthened, video_url is the edit route: the prompt describes the change and the output follows the source clip's framing. It cannot be combined with image_url, end_image_url or the reference fields.
What should I do before October 22?
List which Lite calls use only text and a first or last frame; those port to Omni with a rename and a duration check, since Lite takes 4, 6 or 8 seconds and Sume's Omni takes 3 to 10. Anything that needed references is outside Lite's listed inputs, so check which model it was really calling. Then read capabilities from GET /v1/video-router/models and encode the limits there instead of in prompts.
Sources
Related posts
More in Models
- 9:16 vertical AI video: which Sume models take an aspect ratio
Seedance, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash take 9:16; Grok Imagine, Genjutsu and H3 Max Recast take no aspect ratio. A per-model table.
- AI video ratios on Sume: Omni Flash is 16:9 or 9:16, Kling adds 1:1
Gemini Omni Flash 1.1 on Sume takes only 16:9 and 9:16; Kling 3 adds 1:1. Which other Sume video models list more ratios, and which list none.
- Video-to-video clip length limits on Sume, by tool
Recast takes 5 to 30 s, Genjutsu 4 to 30 s, Omni edit 3 to 10 s, Avatar face swap about 4 to 15 s. A table and a Python picker choose by clip length.
- Vietnamese, Thai, Indonesian, Malay TTS API: vi, th, id, ms on Sume
Cartesia Sonic 3.6 lists vi, th, id and ms. How to send each as the language on Sume TTS, why the field is required, and a per-character cost for each script.
Written by Sume