Vidu @subject tags vs Sume Omni <IMAGE_REF_0> reference tokens
Vidu Q4 reference prompts use @[subject_name]. On Sume only Gemini Omni Flash 1.1 documents prompt tokens: <IMAGE_REF_0>, <VIDEO_REF_0>.

Vidu Q4 Preview's reference-to-video API tells you to point at a reference inside the prompt with @[subject_name]. Sume's catalog documents prompt tokens for one video row only: Gemini Omni Flash 1.1, where references are addressed as <IMAGE_REF_0> and <VIDEO_REF_0>. The index is 0-based and follows the order of your list, and reference media is sent before the prompt. For the other Sume rows we found no documented tag syntax, so describe each reference in plain words.
Side by side
Vidu's page states that the prompt can be up to 20,000 characters and uses the @[subject_name] form to refer to a subject. Sume's Omni constraints also list a prompt limit of 20,000 characters, and one detail matters when you port prompts: a single reference image with no first or end frame is reference-to-video on Sume, and the tag index counts images and videos separately.
| Item | Vidu Q4 Preview | Sume Omni Flash 1.1 |
|---|---|---|
| Tag syntax | @[subject_name] | <IMAGE_REF_0>, <VIDEO_REF_0> |
| Index base | Named subject | 0-based, list order |
| Prompt limit | 20,000 characters | 20,000 characters |
| Max reference images | 15 | 10 |
| Max reference videos | Not listed on the page read | 3, each up to 3 s |
| Audio references | Up to 3 | None |
Port a Vidu prompt to Omni
The mapping is mechanical: replace each named subject with the token for its position in your input_references list.
Steps:
- List your images in a fixed order and write the order down.
- Replace
@[hero]with<IMAGE_REF_0>, the next subject with<IMAGE_REF_1>, and so on. - Keep video references in their own count: the first clip is
<VIDEO_REF_0>. - Send the request with
modelset togemini-omni-flash-1.1and a 16:9 or 9:16 aspect ratio.
{
"model": "gemini-omni-flash-1.1",
"prompt": "<IMAGE_REF_0> hands the bottle to <IMAGE_REF_1> in a sunlit kitchen",
"resolution": "720p",
"duration": 8,
"aspect_ratio": "16:9",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/hero.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/friend.png"}}
]
}Keeping prompts portable
If the same brief will run on several models, write the prompt in two layers: a plain description of each reference that works anywhere, and a tag layer that you add per target. For example, keep "the woman in the red coat" in the base text and let a small script insert the model's tag next to it.
Store the reference list and the prompt in the same record. The risk with 0-based tokens is silent mis-pointing: if someone inserts an image at the front of the list, every tag shifts by one and the clip still renders, just with the wrong person. A test that counts tokens against list length catches most of that.
Keeping references straight
Reference tokens are position-based on Sume, so the order of your reference_image_urls array matters. Keep a small table in your notes that maps each index to its subject, and rebuild the prompt from that table rather than editing by hand. Note that Omni indexes start at zero, so the first image is <IMAGE_REF_0>, which is easy to get wrong when you port a prompt written for one-based tags.
What Sume does not do
Sume does not translate @[name] tags for you; a tag that the model does not know is just text in the prompt. It also does not document a named-subject feature like Vidu's. If the order of your list changes, the tokens point at different images, so keep the list and the prompt together in one file.
Sources
Related posts
More in Developers
- Wait for an avatar video in one request? Sume sync stops at 30 s
Sume sync and subscribe modes are the same bounded wait of at most 30 seconds. For avatar videos use async with polling or a webhook, not a long HTTP hold.
- Wan 3.0 draft and final need different Idempotency-Keys (409)
Reusing one Idempotency-Key for a 480p draft and a 1080p final of the same prompt returns 409 idempotency_conflict. A Python key builder that avoids it.
- wan-3.0 duration 1 gives 400: the 2 to 30 s range, 25 cents at 2 s
Sume lists wan-3.0 at 2 to 30 seconds. A 1-second request is a 400 unsupported_capability; a 2-second 720p clip bills 25 cents (list $0.10 per second x 1.25).
- Wan 3.0 reference limits on Sume: 10 images, 5 videos, 5 audio clips
Wan 3.0 on Sume takes up to 10 reference images, 5 reference videos (15 s total, 16 fps or more) and 5 audio clips (15 s total). Here is the full envelope.
Written by Sume