Omni reference limits: 10 images and 3 clips versus Veo's 3 images

On Sume, Gemini Omni takes up to 10 reference images and 3 clips of 3 seconds each; Veo 3.1 takes three images and Lite none. Build the request in Python.

4 min readSume
All posts

Gemini Omni on Sume accepts up to 10 reference images and up to 3 reference clips of at most 3 seconds each in one request, while Google's Veo 3.1 and Veo 3.1 Fast take up to three reference images and Veo 3.1 Lite takes none. If your Veo workflow ran out of reference slots, this is the practical upgrade path.

Side by side

Google's Veo guide says reference images are allowed up to three on Veo 3.1 and Fast, need an 8-second duration, and are not available on Lite. Google's Omni guide says video references support at most 3 clips of up to 3 seconds each. Sume's Video Router row repeats the Omni envelope and adds the image count.

Reference inputs by model (read 2026-10-05)
ModelReference imagesReference clipsDuration rule
Veo 3.1 / FastUp to 3None listed8 seconds
Veo 3.1 LiteNoneNone4, 6 or 8 seconds
Gemini Omni (Google)Any media allowed3 clips, 3 s eachUp to 10 seconds
gemini-omni-flash-1.1 (Sume)Up to 103 clips, 3 s each3 to 10 seconds

How references enter the request

On /v1/videos, send references in input_references, each entry typed image_url or video_url. A single reference image with no frame field is treated as reference-to-video, so it conditions the whole clip and does not pin the first frame. The Video Router doc says references are addressed in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>. Omni does not accept reference_audio_urls.

A builder that checks the counts

The function below refuses to send more than the documented limits, so a bad request costs you nothing instead of a 400 round trip.

import os, requests

def omni_refs(prompt, images, clips):
    if len(images) > 10 or len(clips) > 3:
        raise ValueError("Omni allows 10 images and 3 clips")
    refs = [{"type": "image_url", "image_url": {"url": u}} for u in images]
    refs += [{"type": "video_url", "video_url": {"url": u}} for u in clips]
    return {"model": "gemini-omni-flash-1.1", "prompt": prompt,
            "duration": 8, "resolution": "720p", "aspect_ratio": "9:16",
            "input_references": refs}

body = omni_refs(
    "The woman in <IMAGE_REF_0> walks the street from <VIDEO_REF_0>, golden light.",
    ["https://example.com/face.jpg"], ["https://example.com/street-3s.mp4"])
r = requests.post("https://api.sume.com/v1/videos", json=body, timeout=60,
                  headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
                           "Idempotency-Key": "omni-refs-001"})
print(r.status_code, r.text[:300])

Picking references well

More slots do not mean better results. A long list of references can pull the clip in several directions, so start with two or three and add more only when a detail keeps drifting.

  • Use one clean, front-facing image per person.
  • Use one image for the setting and one for style, instead of ten near-duplicates.
  • Keep each reference clip to the three seconds that show the motion you want.
  • Name each reference in the prompt, so the model knows which one controls what.

Trim the clips first

Three-second clips are short. If your source is longer, cut it with video trim, which costs $0.02 a job. Every reference URL must be reachable over public HTTPS, and the Videos guide says a failed download shows up as a failed job with a message that names the field, not the URL.

Sources

Related posts

More in Models

All Models posts

Written by Sume