Veo 3.1 extend adds 7 s up to 20 times, 720p only: plan a long clip
Google's Veo docs let you extend a clip by 7 seconds up to 20 times, at 720p only. Read the arithmetic, then compare it with one-call 30 second models on Sume.

Google's Veo guide says a Veo 3.1 clip can be extended by 7 seconds, up to 20 times, and that extension works at 720p only. If you start from an 8 second base, the two figures give a ceiling of 148 seconds, with every second after the base at 720p.
The 148 is derived arithmetic from those figures, not a number Google states. The point of the post is to show what the limit means for a long ad, and what to compare it with.
The Veo facts, as read on 2026-10-03
These come from the Gemini API Veo page. Veo durations are 4, 6 or 8 seconds, and an 8 second duration is needed for 1080p, 4K and reference images.
| Item | Value |
|---|---|
| Base durations | 4, 6 or 8 seconds |
| Needs an 8 second duration | 1080p, 4K and reference images |
| Extension step | 7 seconds |
| Extensions allowed | Up to 20 |
| Extension resolution | 720p only |
| Storage of generated videos | 2 days |
What the 720p-only rule changes
A 1080p base clip needs an 8 second duration, but extending it is a 720p operation. That means a chain that starts sharp and grows past the base ends up as a 720p clip, or a clip whose resolution changes along the way. For a delivery spec that demands 1080p, extension is not a way to get a longer 1080p clip.
Because the extended segments are generated from the previous clip, a mistake in an early segment is carried forward. Review the base clip before you spend on extensions.
The same length with one call on Sume
Sume's video docs say limits are not uniform across the catalog. seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p, wan-3.0 accepts 2 to 30 seconds, and every other catalog model tops out at 15 seconds. Gemini Omni Flash 1.1 accepts 3 to 10 seconds. So a 30 second shot can be one request on a model that supports it, instead of a base plus extensions.
For anything longer, Timeline 1.0 assembles an audio spine and an ordered video[] list into one MP4; the render takes audio.duration_seconds from 1 to 1800 and 1 to 200 video slots. Read supported_durations in the live catalog before you pick a model, because the docs say to check them.
| Longest single clip | Clips for 60 seconds | Example from the docs |
|---|---|---|
| 30 seconds | 2 | seedance-2.5, wan-3.0 |
| 15 seconds | 4 | Other catalog models |
| 10 seconds | 6 | Gemini Omni Flash 1.1 |
A small planner
This counts how many clips a target length needs for a given maximum per clip. It does no network calls, so it runs as is.
import math
def clips_needed(target_seconds, max_clip_seconds):
if target_seconds <= 0 or max_clip_seconds <= 0:
raise ValueError("both values must be positive")
return math.ceil(target_seconds / max_clip_seconds)
if __name__ == "__main__":
for cap in (30, 15, 10):
print(cap, clips_needed(60, cap))Checklist
- Decide the delivery resolution first; if it is 1080p, do not plan on Veo extension.
- Download Veo outputs inside the 2 day window if you use Veo directly.
- On Sume, read
supported_durationsfrom the catalog for the model you intend to pin. - Join clips with Timeline 1.0 rather than chaining by hand when the cut points are yours.
What this post does not claim
It does not claim that Veo extension is bad. Extension builds on an existing clip, which suits one continuous camera move. It also does not claim that Sume offers Veo extension; the docs read for this post describe single-request durations per model and Timeline assembly, not an extend operation.
If continuity of a single shot matters more than resolution, use Google's extension and accept 720p. If resolution matters more, plan separate shots at the target resolution and cut between them, using a last frame as the next first frame where the model accepts a first-frame image.
Sources
Related posts
More in Models
- Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.
- Wan 3.0 doubles clip length to 30 seconds: Alibaba ids vs Sume
Alibaba Model Studio lists wan3.0-video and wan3.0-video-prime at 2 to 30 seconds, up from 15 on Wan 2.7. What Sume's wan-3.0 accepts and a 30 s request.
- What is FLUX 3? The family map: image, video, audio, action
BFL's docs describe FLUX 3 as one family covering image, video with synchronized audio, audio and action. Which pieces have open weights and which do not.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume