Gemini Omni [# Sources] and [# References] tags vs Sume fields
Gemini Omni binds media to roles with tags like <FIRST_FRAME> and [# Sources ...]. Sume uses request fields instead: image_url, end_image_url, reference lists.

In Gemini Omni Flash prompts, tags bind uploaded media to a role: <FIRST_FRAME>, <LAST_FRAME>, <IMAGE_REF_N> and <VIDEO_REF_N>, with a longer form that declares media up front as [# Sources ...] and [# References ...]. On Sume you do not write the frame tags; you send image_url for the first frame and end_image_url for the last, and only the references are addressed inside the prompt, as <IMAGE_REF_0> and <VIDEO_REF_0>.
Google's side is from its Gemini API: Generate and edit videos with Gemini Omni Flash, Sume's from Video Router, both read 2026-10-02. Sume's page does not describe the [# Sources ...] declaration or the frame tags, so this post treats them as Google-only.
What do the simple tags do?
Google calls the simple tags the recommended form for cases where roles are clear. <FIRST_FRAME> uses an image as the starting frame; <LAST_FRAME> is the final frame and must be used together with <FIRST_FRAME>. <IMAGE_REF_N> and <VIDEO_REF_N> mark references, and both count from 0.
Google's example combines a style reference and a subject reference in one sentence: in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking.
What is the [# Sources] form for?
For many media inputs with several roles, Google says to declare them at the start of the prompt, for example [# Sources <FIRST_FRAME>@Image1 <LAST_FRAME>@Image2] or [# References <IMAGE_REF_0>@Image1 <VIDEO_REF_0>@Video1]. It also lists <VIDEO_0> for the source video to edit and <PREVIOUS_VIDEO> for the clip to extend, and recommends closing the prompt with a guiding sentence such as "Use the given image(s) as references for video generation."
Google's page shows using the same image as both <FIRST_FRAME> and <LAST_FRAME> to make a loop.
How do the same roles map onto Sume's request?
The Video Router page says one catalog id routes by the shape of the request: prompt alone is text to video; image_url with an optional end_image_url is image to video; reference_image_urls (up to 10) and reference_video_urls (up to 3, each at most 3 seconds) are reference to video; video_url is an edit. The video_url field cannot be combined with image_url, end_image_url or the reference lists.
| Role | Google prompt tag | Sume field |
|---|---|---|
| First frame | <FIRST_FRAME> | image_url |
| Last frame | <LAST_FRAME> (needs first frame) | end_image_url (needs image_url) |
| Image reference | <IMAGE_REF_N> | reference_image_urls, addressed as <IMAGE_REF_0> |
| Video reference | <VIDEO_REF_N> | reference_video_urls, addressed as <VIDEO_REF_0> |
| Clip to edit | <VIDEO_0> in [# Sources ...] | video_url |
What should I do when porting a Google prompt to Sume?
Three changes cover most prompts.
- Move first-frame and last-frame images out of the prompt and into
image_urlandend_image_url. - Keep
<IMAGE_REF_N>and<VIDEO_REF_N>in the prompt; keep indexes 0-based in list order. - Do not send
video_urltogether with frames or references; split it into two jobs.
Sources
Related posts
More in Models
- Gemini Omni and Veo in Korean: only English is fully supported
Google's Omni and Veo pages say English is fully supported; other languages aren't evaluated. For Korean prompts, describe in English and quote on-screen text.
- Gemini Omni can't use a YouTube link as its source video
Google lists YouTube videos as an unsupported media source for Gemini Omni Flash. On Sume, video_url is a URI field; use a direct link to the clip file.
- GPT Image 2.5 mask_url edit: change one region, keep the rest
GPT Image 2.5 on Sume accepts a public mask_url alongside input_references. Build a mask with Pillow, host it, and edit only one region of a photo.
- GPT Image 2.5 character drift: what OpenAI says and what to send
OpenAI's image guide says GPT Image may struggle to keep a recurring character the same across generations. What to send on Sume to reduce drift.
Written by Sume