Veo 3.1 reference images and extension: Google yes, Sume text-only
Google's Veo 3.1 takes up to 3 reference images and extends clips 7 seconds at a time. Sume's Veo rows are text-to-video only. What to use for references.

On Google's Gemini API, Veo 3.1 accepts up to 3 reference images and can extend a Veo clip by 7 seconds, up to 20 times (read 2026-10-10). Sume's Veo rows do neither: in the API code they are text-to-video only and refuse frame, reference and video inputs.
If your Veo plan depends on references or extension, Sume's Veo rows will not do the job. Use a different Sume row for references, and keep extension work on Google.
What Google's page says
The Gemini API Veo page (read 2026-10-10) lists three ids: veo-3.1-generate-preview, veo-3.1-fast-generate-preview and veo-3.1-lite-generate-preview. Extension works on 3.1 and Fast only, adds 7 seconds each time, and can repeat up to 20 times. Reference images cap at 3.
Google stores generated videos for 2 days on that page, so download or extend inside the window.
| Input | Google Gemini API | Sume Veo rows |
|---|---|---|
| Text prompt | Yes | Yes |
| Start or end frame | Not covered here | Refused |
| Reference images | Up to 3 | Refused |
| Extend a clip | 3.1 and Fast, 7 s each | Refused |
How Sume's Veo rows behave
Sume's request builder for Veo (buildVeoRequest in the API code) takes the prompt, a duration of 4, 6 or 8 seconds, 720p, an aspect ratio of 16:9 or 9:16, and an optional audio flag. It throws on frame fields, reference fields and video fields. The catalog description reads "text-to-video only, 720p, 4/6/8 seconds, 16:9 or 9:16, with optional synchronized audio."
These rows are switched on per environment. They show up in your workspace only if GET /v1/videos/models returns the id, so check the catalog and do not assume either way. An earlier post, linked below, covers the same availability question.
What to use for references on Sume
The Video Generation docs say Seedance 2.x, Wan 3.0 and the MiniMax H3 rows accept image, video and audio references, and Gemini Omni Flash 1.1 accepts image and video references. In the schema, Wan 3.0 takes up to 10 images, 5 videos and 5 audio files. Omni takes up to 10 images and 3 reference videos of at most 3 seconds each.
None of these is Veo, so look at your reference clip first. If you need Google's look, stay on Google. If you need consistency with a character sheet, a reference-capable row is the practical stand-in.
- References wanted: Wan 3.0, Seedance 2.5 or Omni.
- Extension wanted: Seedance 2.5 or Wan 3.0 reach 30 s in one request.
- Veo look wanted: use Google's API.
A longer shot without extension
Extension at Google adds 7 seconds per step. On Sume you get length in one request instead: Wan 3.0 and Seedance 2.5 both accept up to 30 seconds. The table compares an 8-second Veo clip with a 30-second shot on those rows, using Sume's billed rates (list times 1.25) and the Veo row only if it is listed in your workspace.
Cost is not the only difference. A single 30-second request has no seams, while a chain of 7-second extensions inherits the previous clip's last frames. But you also lose Veo's look, so treat this as a substitute, not a swap.
| Sume id | Clip | Billed cost |
|---|---|---|
| veo-3.1 (if listed) | 8 s, 720p, audio | $4.00 |
| veo-3.1-fast (if listed) | 8 s, 720p, audio | $1.00 |
| wan-3.0 | 30 s, 720p | $3.75 |
| seedance-2.5 | 30 s, 720p | $17.334 |
Checking support in the catalog
Reference-heavy work is where the choice matters most. Google caps references at 3 images and gives you extension; Sume's reference-capable rows allow more inputs but no extension step. If the 3-image limit was your constraint, Wan 3.0's 10 images and Omni's 10 images are a real gain, and Seedance, Wan and MiniMax also take audio references that Veo's page does not list.
Whatever row you choose, read supported_input_references and supported_frame_images from GET /v1/videos/models before you send reference files. A model accepts a reference type only if the list includes it, per the Video Generation docs.
Sources
Related posts
More in Comparisons
- Vidu Q4 clip lengths 3 to 16 s vs every Sume video row's limits
Vidu's image-to-video page allows 3-16 seconds. Sume has no Vidu row; this table maps that range to each Sume row's minimum and maximum, and who reaches 16.
- Vidu Q4 has no text-to-video: Sume rows that take a plain prompt
Vidu's Q4 page lists only image-to-video and reference-to-video. Sume has no Vidu row, but several rows start from a text prompt alone. Compare limits.
- Vidu Q4 takes 15 reference images; Sume rows cap at 9 or 10
Vidu Q4 accepts 1 to 15 reference images. Sume has no Vidu row: most rows allow 9 images, Wan 3.0 and Omni allow 10. Caps for images, video and audio refs.
- Vidu Q4 voice references vs Sume's reference audio inputs
Vidu Q4 takes up to 3 voice references. Sume has no Vidu row; Seedance, Wan 3.0 and MiniMax H3 accept audio references, but voice cloning is app-only, not API.
Written by Sume