<IMAGE_REF_0> and <VIDEO_REF_0>: tag references in Gemini Omni

On gemini-omni-flash-1.1 you point at each reference by token in the prompt: <IMAGE_REF_0>, <VIDEO_REF_0>, numbered from 0 in list order. Request example.

4 min readSume
All posts

In a reference-to-video request to gemini-omni-flash-1.1 on Sume, you refer to each reference file by a token inside the prompt: <IMAGE_REF_0> for the first reference image and <VIDEO_REF_0> for the first reference video. Numbering is 0-based and follows the order of your list. The catalog notes add that reference media is sent before the prompt.

The request shape

Reference-to-video on this model uses reference_image_urls (up to 10) and reference_video_urls (up to 3, each at most 3 seconds). There is no audio reference. A single reference image with no first or end frame is treated as reference-to-video by the catalog, so you do not need a mode flag. Frames and references cannot be mixed, as the Video Router refuses both in one request.

Gemini Omni Flash 1.1 reference rules in the Sume catalog and docs on main (read 2026-10-05)
ItemRule
Reference imagesUp to 10, tokens <IMAGE_REF_0>, <IMAGE_REF_1>, ... in list order
Reference videosUp to 3, each at most 3 s, tokens <VIDEO_REF_0>, ...
Reference audioNot supported
Duration3 to 10 s
Aspect ratio16:9 or 9:16
Audio outputNative synced audio, always on

Using the tokens

Write the prompt so each token stands where you would otherwise write a description of that file. For example, "<IMAGE_REF_0> walks into the room from <IMAGE_REF_1>, moving like <VIDEO_REF_0>." The tokens make it unambiguous which picture is the person and which is the place, instead of leaving the model to guess from the order in a sentence. If you reorder the list, renumber the tokens.

A complete request

The body below is a valid shape for the Video Router. It only sends when you set SEND=1, since a valid request is a billed job.

import os
import requests

body = {
    "model": "gemini-omni-flash-1.1",
    "prompt": "<IMAGE_REF_0> walks into the room from <IMAGE_REF_1>, moving like <VIDEO_REF_0>.",
    "reference_image_urls": ["https://example.com/person.png", "https://example.com/room.png"],
    "reference_video_urls": ["https://example.com/walk-3s.mp4"],
    "resolution": "720p",
    "duration": 6,
    "aspect_ratio": "16:9",
}
print(sorted(body))
if os.environ.get("SEND") == "1":
    r = requests.post("https://api.sume.com/v1/video-router/generate",
                      headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
                      json=body, timeout=60)
    print(r.status_code)

When tokens are the wrong tool

Tokens are for reference-to-video. In edit mode the source is a video_url and the prompt describes the change in plain words; references and tokens do not apply there. If you need a specific person from a photo in an edit, that is not available in edit mode, as the related post on Gemini Omni edit explains. Keep reference videos short: the three-second limit is the first thing a longer clip will trip.

A few habits make token prompts reliable. Name each token once and describe the action, not the appearance, since the reference already carries the appearance. Keep the list short: the model can take 10 images, but a prompt that refers to eight different tokens is hard to read for people and models alike. And log the final prompt with the list order, so a changed result can be traced to a changed order rather than a changed model.

Sources

Related posts

More in Models

All Models posts

Written by Sume