Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.

To make three different characters perform three different dances in one clip, send three images as reference_image_urls and three short clips as reference_video_urls to gemini-omni-flash-1.1, then pair them in the prompt with <IMAGE_REF_0> to <IMAGE_REF_2> and <VIDEO_REF_0> to <VIDEO_REF_2>. The tokens are 0-based and follow the order of each list. Each video reference can be at most 3 seconds.
What Google demonstrates
The launch post's example prompt takes three uploaded dance videos and replaces the dancers with provided characters: a dog does a classical dance, an octopus a hip hop dance, a bear a breakdance, all together in a large open space from a provided image, as one continuous shot with no cuts. It also says the model can reference up to three seconds of video to keep visual context and character consistency.
The same request on Sume
Sume lists reference-to-video as reference_image_urls (up to 10) and/or reference_video_urls (up to 3, each up to 3 seconds), with no audio references. The setting for the room is a fourth image, so it takes index 3.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: three-dancers-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Place <IMAGE_REF_0>, <IMAGE_REF_1> and <IMAGE_REF_2> together in the room from <IMAGE_REF_3>. <IMAGE_REF_0> performs the dance from <VIDEO_REF_0>, <IMAGE_REF_1> performs the dance from <VIDEO_REF_1>, <IMAGE_REF_2> performs the dance from <VIDEO_REF_2>. One continuous shot, no scene cuts.",
"reference_image_urls": ["https://example.com/dog.png","https://example.com/octopus.png","https://example.com/bear.png","https://example.com/room.png"],
"reference_video_urls": ["https://example.com/dance1.mp4","https://example.com/dance2.mp4","https://example.com/dance3.mp4"],
"duration": 8,
"resolution": "720p",
"mode": "async"
}'Order table
Keep a small table next to your request. Reordering a list silently reassigns every token that comes after the change.
| Token | List position | File |
|---|---|---|
| <IMAGE_REF_0> | reference_image_urls[0] | dog.png |
| <IMAGE_REF_1> | reference_image_urls[1] | octopus.png |
| <IMAGE_REF_2> | reference_image_urls[2] | bear.png |
| <IMAGE_REF_3> | reference_image_urls[3] | room.png |
| <VIDEO_REF_0> | reference_video_urls[0] | dance1.mp4 |
| <VIDEO_REF_1> | reference_video_urls[1] | dance2.mp4 |
| <VIDEO_REF_2> | reference_video_urls[2] | dance3.mp4 |
Writing the prompt around the tokens
Say each pairing once and explicitly. A prompt that only lists the characters and the dances leaves the model to guess which move belongs to whom. Name the shot as continuous, name the space, and say what should not happen, such as no cuts. Google's example does the same: it gives the setting, the three pairings and the single-shot instruction in one paragraph.
Pick dance clips that are visually distinct, with the full body in frame, and keep each reference to its three seconds of cleanest movement.
Native audio is generated with the clip and cannot be switched off on this model, and there is no audio reference input, so the dance clips' own soundtracks are not carried over. If the track matters, lay the music over the result afterwards.
Limits and cost
The output is 3 to 10 seconds, so a dance that needs 15 seconds has to be cut into two requests. An 8-second clip at 720p is $1.00 on Sume ($0.125 per second, list x 1.25), and $1.50 at 1080p. If a reference clip runs over 3 seconds, trim it first with video trim at $0.02 per job, and see the reference error cases for what a 4th clip or a 4-second clip returns.
For one dance move on one character with a longer reference, Kling motion control takes a motion video of 1 to 30 seconds instead.
Sources
Related posts
More in Models
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
- Minimum clip length by Sume video model: 2, 3, 4 or 5 seconds
Wan 3.0 starts at 2 s, Gemini Omni Flash 1.1 at 3 s, Seedance and Genjutsu at 4 s, MiniMax H3 and Recast at 5 s. Full min and max table for a Sora port.
- Music Router image_url: public HTTPS only, and null to clear it
Sume music image_url must be a public HTTPS image, and null clears it on a reused request object. When a mood still helps and when the prompt does more.
- Nano Banana 2 is retired on Sume: same price on Nano Banana 2.1?
Sume retired google/nano-banana-2 and runs old requests as Nano Banana 2.1 at the same price card: $0.10 at 1K billed. Tiers, ids and what the job stores.
Written by Sume