Omni reference limits: 10 images and 3 clips versus Veo's 3 images
On Sume, Gemini Omni takes up to 10 reference images and 3 clips of 3 seconds each; Veo 3.1 takes three images and Lite none. Build the request in Python.

Gemini Omni on Sume accepts up to 10 reference images and up to 3 reference clips of at most 3 seconds each in one request, while Google's Veo 3.1 and Veo 3.1 Fast take up to three reference images and Veo 3.1 Lite takes none. If your Veo workflow ran out of reference slots, this is the practical upgrade path.
Side by side
Google's Veo guide says reference images are allowed up to three on Veo 3.1 and Fast, need an 8-second duration, and are not available on Lite. Google's Omni guide says video references support at most 3 clips of up to 3 seconds each. Sume's Video Router row repeats the Omni envelope and adds the image count.
| Model | Reference images | Reference clips | Duration rule |
|---|---|---|---|
| Veo 3.1 / Fast | Up to 3 | None listed | 8 seconds |
| Veo 3.1 Lite | None | None | 4, 6 or 8 seconds |
| Gemini Omni (Google) | Any media allowed | 3 clips, 3 s each | Up to 10 seconds |
| gemini-omni-flash-1.1 (Sume) | Up to 10 | 3 clips, 3 s each | 3 to 10 seconds |
How references enter the request
On /v1/videos, send references in input_references, each entry typed image_url or video_url. A single reference image with no frame field is treated as reference-to-video, so it conditions the whole clip and does not pin the first frame. The Video Router doc says references are addressed in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>. Omni does not accept reference_audio_urls.
A builder that checks the counts
The function below refuses to send more than the documented limits, so a bad request costs you nothing instead of a 400 round trip.
import os, requests
def omni_refs(prompt, images, clips):
if len(images) > 10 or len(clips) > 3:
raise ValueError("Omni allows 10 images and 3 clips")
refs = [{"type": "image_url", "image_url": {"url": u}} for u in images]
refs += [{"type": "video_url", "video_url": {"url": u}} for u in clips]
return {"model": "gemini-omni-flash-1.1", "prompt": prompt,
"duration": 8, "resolution": "720p", "aspect_ratio": "9:16",
"input_references": refs}
body = omni_refs(
"The woman in <IMAGE_REF_0> walks the street from <VIDEO_REF_0>, golden light.",
["https://example.com/face.jpg"], ["https://example.com/street-3s.mp4"])
r = requests.post("https://api.sume.com/v1/videos", json=body, timeout=60,
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "omni-refs-001"})
print(r.status_code, r.text[:300])
Picking references well
More slots do not mean better results. A long list of references can pull the clip in several directions, so start with two or three and add more only when a detail keeps drifting.
- Use one clean, front-facing image per person.
- Use one image for the setting and one for style, instead of ten near-duplicates.
- Keep each reference clip to the three seconds that show the motion you want.
- Name each reference in the prompt, so the model knows which one controls what.
Trim the clips first
Three-second clips are short. If your source is longer, cut it with video trim, which costs $0.02 a job. Every reference URL must be reachable over public HTTPS, and the Videos guide says a failed download shows up as a failed job with a message that names the field, not the URL.
Sources
Related posts
More in Models
- GPT Image 2.5 can take 2 minutes: submit async, not sync, on Sume
OpenAI says complex GPT Image 2.5 prompts can run up to 2 minutes. Sume's image route blocks only 30 seconds, so send mode async and poll the job.
- GPT Image 2.5 edits through Sume: mask_url and 16 references
Sume's GPT Image 2.5 takes up to 16 references, an optional mask_url and background auto, transparent or opaque. OpenAI: transparent needs png or webp.
- Pocket TTS license: the repo says MIT, not Apache-2.0
Some roundups call Kyutai's Pocket TTS Apache-2.0. Its GitHub page and LICENSE file read MIT-style. How to check a TTS license before you build on it.
- Pocket TTS runs on 2 CPU cores: what ~200 ms first audio means
Kyutai's Pocket TTS lists 100M parameters, 2 CPU cores and ~200 ms to first audio. Whether that matters for a video voiceover, and a hosted TTS job's numbers.
Written by Sume