Two reference images, one clip: Omni's cat-and-yarn pattern on Sume

Google's docs show two images, a cat and yarn, producing one video. Here is the Sume request with IMAGE_REF tokens and what changes with a brand product.

4 min readSume
All posts

To put two reference subjects in one Gemini Omni Flash 1.1 clip, send both pictures in reference_image_urls and refer to them in the prompt as <IMAGE_REF_0> and <IMAGE_REF_1>. Google's documentation uses a cat and a ball of yarn as the example. The tokens are 0-based, and up to 10 images are allowed on Sume.

What Google shows

The Gemini API docs describe subject reference as generating a video that incorporates specific subjects provided as reference images, and show two PNGs, a cat and yarn, with the prompt "A cat playfully batting at a ball of yarn." (read 2026-10-07). The Google request passes images as input parts and writes no tokens. Sume's reference mode adds the tokens, so the prompt can say which image plays which role.

The Sume request

Reference mode uses reference_image_urls and/or reference_video_urls, with no image_url or video_url in the same request. The prompt names each image by its index.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cat-yarn-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "<IMAGE_REF_0> playfully bats at <IMAGE_REF_1> on a wooden floor. Handheld, warm afternoon light.",
    "reference_image_urls": [
      "https://example.com/cat.png",
      "https://example.com/yarn.png"
    ],
    "duration": 6,
    "resolution": "720p",
    "mode": "async"
  }'

Swap in your own subject and prop

The pattern is subject plus object, so it carries over to a person and a product, a mascot and a prop, or a pet and a toy. Use one clean image per subject, on a plain or neutral background, so the model does not carry the old background into the clip. State the interaction in the prompt in one verb phrase; the cat example is "playfully batting", and a clear verb gives the model something to animate.

Two-image reference request, 6 seconds, Sume (read 2026-10-07)
ResolutionPer second6 seconds
360p$0.0375$0.225
720p$0.125$0.75
1080p$0.1875$1.125

Camera and motion

Reference mode is subject-driven, so the prompt carries the camera. Say "handheld", "static" or "slow push-in" once, and say where the action happens. If you leave the setting out, the model fills it from the images, which can pull the background of one picture into the whole clip. A neutral location named in the prompt, such as a wooden floor in afternoon light, keeps the reference backgrounds from leaking in.

Six seconds is long enough for one beat of action. A cat bats at yarn, the yarn rolls, the cat follows. Asking for a story with three events in 6 seconds produces rushed motion.

When it goes wrong

If the prompt says nothing about which image is which, the model may merge them. If a token points past the end of the list, the request is a mistake in your payload, not a creative problem; keep the list and the prompt in the same file so you edit them together. More than 10 images is rejected, and so is more than 3 reference videos; the reference error cases list the messages. Native audio is on and cannot be turned off, so even a silent-looking subject gets sound. Run a 360p draft first at $0.225 for 6 seconds to see whether the two subjects read as separate things.

To reuse the same character across a whole sequence, see reference clips of three seconds each.

Sources

Related posts

More in Models

All Models posts

Written by Sume