HappyHorse 1.1 [Image 1] tags vs Sume's <IMAGE_REF_0> in prompts

HappyHorse reference-to-video counts images from [Image 1]; Sume's Omni counts from <IMAGE_REF_0>. How to map a prompt without an off-by-one.

4 min readSume
All posts

Alibaba's HappyHorse reference-to-video reference says to refer to images in the prompt as [Image 1], [Image 2], by position, and accepts 1 to 9 reference images (Alibaba Model Studio, read 2026-10-09). Sume's gemini-omni-flash-1.1 takes up to 10 images and uses <IMAGE_REF_0>, counting from zero. A prompt ported tag for tag is off by one.

How do the two tag systems map?

Prompt tags (read 2026-10-09)
Position in your listHappyHorse 1.1 r2vSume Omni
First image[Image 1]<IMAGE_REF_0>
Second image[Image 2]<IMAGE_REF_1>
Third image[Image 3]<IMAGE_REF_2>
First reference clipNot listed<VIDEO_REF_0>

How do you port the prompt safely?

Write the references as a list in code, build the tags from the index, and let the code subtract one. A short script does it:

import re

def to_sume(prompt: str) -> str:
    return re.sub(r"\[Image (\d+)\]", lambda m: f"<IMAGE_REF_{int(m.group(1)) - 1}>", prompt)

print(to_sume("the woman in [Image 1] hands a cup to [Image 2]"))
# the woman in <IMAGE_REF_0> hands a cup to <IMAGE_REF_1>

What else differs?

Sume lists no HappyHorse id (catalog read 2026-10-09), so this is a prompt migration, not a model swap. The Video Router docs state the tag syntax.

  • Image size: HappyHorse wants the shortest side at least 400 px and at most 20 MB.
  • Count: 1 to 9 images on HappyHorse; up to 10 images and 3 clips of 3 seconds on Sume.
  • Length: 3 to 15 seconds on HappyHorse; 3 to 10 on Sume's Omni row.
  • Ratios: HappyHorse lists nine; Omni on Sume lists 16:9 and 9:16.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume