Gemini Omni [# Sources] and [# References] tags vs Sume fields

Gemini Omni binds media to roles with tags like <FIRST_FRAME> and [# Sources ...]. Sume uses request fields instead: image_url, end_image_url, reference lists.

5 min readSume
All posts

In Gemini Omni Flash prompts, tags bind uploaded media to a role: <FIRST_FRAME>, <LAST_FRAME>, <IMAGE_REF_N> and <VIDEO_REF_N>, with a longer form that declares media up front as [# Sources ...] and [# References ...]. On Sume you do not write the frame tags; you send image_url for the first frame and end_image_url for the last, and only the references are addressed inside the prompt, as <IMAGE_REF_0> and <VIDEO_REF_0>.

Google's side is from its Gemini API: Generate and edit videos with Gemini Omni Flash, Sume's from Video Router, both read 2026-10-02. Sume's page does not describe the [# Sources ...] declaration or the frame tags, so this post treats them as Google-only.

What do the simple tags do?

Google calls the simple tags the recommended form for cases where roles are clear. <FIRST_FRAME> uses an image as the starting frame; <LAST_FRAME> is the final frame and must be used together with <FIRST_FRAME>. <IMAGE_REF_N> and <VIDEO_REF_N> mark references, and both count from 0.

Google's example combines a style reference and a subject reference in one sentence: in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking.

What is the [# Sources] form for?

For many media inputs with several roles, Google says to declare them at the start of the prompt, for example [# Sources <FIRST_FRAME>@Image1 <LAST_FRAME>@Image2] or [# References <IMAGE_REF_0>@Image1 <VIDEO_REF_0>@Video1]. It also lists <VIDEO_0> for the source video to edit and <PREVIOUS_VIDEO> for the clip to extend, and recommends closing the prompt with a guiding sentence such as "Use the given image(s) as references for video generation."

Google's page shows using the same image as both <FIRST_FRAME> and <LAST_FRAME> to make a loop.

How do the same roles map onto Sume's request?

The Video Router page says one catalog id routes by the shape of the request: prompt alone is text to video; image_url with an optional end_image_url is image to video; reference_image_urls (up to 10) and reference_video_urls (up to 3, each at most 3 seconds) are reference to video; video_url is an edit. The video_url field cannot be combined with image_url, end_image_url or the reference lists.

Google's Omni guide and Sume's Video Router page, read 2026-10-02.
RoleGoogle prompt tagSume field
First frame<FIRST_FRAME>image_url
Last frame<LAST_FRAME> (needs first frame)end_image_url (needs image_url)
Image reference<IMAGE_REF_N>reference_image_urls, addressed as <IMAGE_REF_0>
Video reference<VIDEO_REF_N>reference_video_urls, addressed as <VIDEO_REF_0>
Clip to edit<VIDEO_0> in [# Sources ...]video_url

What should I do when porting a Google prompt to Sume?

Three changes cover most prompts.

  • Move first-frame and last-frame images out of the prompt and into image_url and end_image_url.
  • Keep <IMAGE_REF_N> and <VIDEO_REF_N> in the prompt; keep indexes 0-based in list order.
  • Do not send video_url together with frames or references; split it into two jobs.

Sources

Related posts

More in Models

All Models posts

Written by Sume