<IMAGE_REF_0> and <VIDEO_REF_0>: tag references in Gemini Omni
On gemini-omni-flash-1.1 you point at each reference by token in the prompt: <IMAGE_REF_0>, <VIDEO_REF_0>, numbered from 0 in list order. Request example.

In a reference-to-video request to gemini-omni-flash-1.1 on Sume, you refer to each reference file by a token inside the prompt: <IMAGE_REF_0> for the first reference image and <VIDEO_REF_0> for the first reference video. Numbering is 0-based and follows the order of your list. The catalog notes add that reference media is sent before the prompt.
The request shape
Reference-to-video on this model uses reference_image_urls (up to 10) and reference_video_urls (up to 3, each at most 3 seconds). There is no audio reference. A single reference image with no first or end frame is treated as reference-to-video by the catalog, so you do not need a mode flag. Frames and references cannot be mixed, as the Video Router refuses both in one request.
| Item | Rule |
|---|---|
| Reference images | Up to 10, tokens <IMAGE_REF_0>, <IMAGE_REF_1>, ... in list order |
| Reference videos | Up to 3, each at most 3 s, tokens <VIDEO_REF_0>, ... |
| Reference audio | Not supported |
| Duration | 3 to 10 s |
| Aspect ratio | 16:9 or 9:16 |
| Audio output | Native synced audio, always on |
Using the tokens
Write the prompt so each token stands where you would otherwise write a description of that file. For example, "<IMAGE_REF_0> walks into the room from <IMAGE_REF_1>, moving like <VIDEO_REF_0>." The tokens make it unambiguous which picture is the person and which is the place, instead of leaving the model to guess from the order in a sentence. If you reorder the list, renumber the tokens.
A complete request
The body below is a valid shape for the Video Router. It only sends when you set SEND=1, since a valid request is a billed job.
import os
import requests
body = {
"model": "gemini-omni-flash-1.1",
"prompt": "<IMAGE_REF_0> walks into the room from <IMAGE_REF_1>, moving like <VIDEO_REF_0>.",
"reference_image_urls": ["https://example.com/person.png", "https://example.com/room.png"],
"reference_video_urls": ["https://example.com/walk-3s.mp4"],
"resolution": "720p",
"duration": 6,
"aspect_ratio": "16:9",
}
print(sorted(body))
if os.environ.get("SEND") == "1":
r = requests.post("https://api.sume.com/v1/video-router/generate",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json=body, timeout=60)
print(r.status_code)
When tokens are the wrong tool
Tokens are for reference-to-video. In edit mode the source is a video_url and the prompt describes the change in plain words; references and tokens do not apply there. If you need a specific person from a photo in an edit, that is not available in edit mode, as the related post on Gemini Omni edit explains. Keep reference videos short: the three-second limit is the first thing a longer clip will trip.
A few habits make token prompts reliable. Name each token once and describe the action, not the appearance, since the reference already carries the appearance. Keep the list short: the model can take 10 images, but a prompt that refers to eight different tokens is hard to read for people and models alike. And log the final prompt with the list order, so a changed result can be traced to a changed order rather than a changed model.
Sources
Related posts
More in Models
- Is Ideogram 4.5 edit a separate id on Sume? No, references decide
No. ideogram/ideogram-v4.5 is the only id: no input_references generates, 1 to 5 edits the first image. Size and aspect rules change with the call type.
- Is the Seedance 2.5 API live? What Sume can call for 30-second video
ByteDance's Seedance 2.5 post said its API was coming soon via BytePlus ModelArk. On Sume, seedance-2.5 (4-30 s) and wan-3.0 (2-30 s) are in the Video Router.
- Japanese and Tagalog TTS: not on MAI-Voice-2.1, tagged on Sume
MAI-Voice-2.1's 23 languages exclude Japanese and Tagalog; Sume's voice library tags ja and tl. How to set language and audition for 1 cent.
- Kling 3.0 native 4K at 60 fps: what Sume's kling-3 lets you order
Kling 3.0 is marketed with native 4K. The fal page lists up to 1080p, and Sume's kling-3 accepts 720p or 1080p, 4 to 15 s. Here is what to plan around.
Written by Sume