Person holding your product: reference clip caps on Omni, Wan and H3

To put a person with your product in one clip, send a photo plus a short reference clip. Omni takes 3 clips of 3 s, Wan 5 clips and 15 s, H3 3 clips and 15 s.

5 min readSume
All posts

To show a person holding your product, send the product photo and a short clip of the person as references. gemini-omni-flash-1.1 takes up to 3 reference clips of 3 seconds or less each; wan-3.0 takes up to 5 clips with 15 seconds in total; minimax-h3 and minimax-h3-max take up to 3 clips with 15 seconds in total. kling-3 and Grok take no references at all.

If you already have the person on film and want to replace them, that is a different tool: h3-max-recast swaps the people in one source video for 1 to 4 photos. Limits here are from the Video Router guide, read 2026-10-05.

Clip caps by model

Wan and H3 also list a minimum clip length and a frame-rate floor; check them against your footage before you upload.

Reference clip limits on Sume (read 2026-10-05)
Model idReference clipsPer-clip or total limitNotes
gemini-omni-flash-1.1up to 3each 3 s or less10 images, no audio references
wan-3.0up to 515 s total, 16 fps or higherup to 10 images and 5 audio clips
minimax-h3up to 32 to 15 s each, 15 s combinedimages, clips and audio together at most 12
minimax-h3-maxup to 32 to 15 s each, 15 s combinedimages, clips and audio together at most 12

Writing the prompt

On Omni, refer to the media as <IMAGE_REF_0> and <VIDEO_REF_0>, numbered from 0 in list order, and write who does what: the person in the clip picks up the item in the photo. The references are sent before the prompt, so the order of the list is the numbering.

Keep the reference clip to the part that shows the person clearly, a few seconds at most on Omni. A 3-second clip with one face and steady light is a better reference than a 15-second crowd scene.

  • One person per reference clip.
  • One product photo with the label facing the camera.
  • Keep the prompt to one action, such as lifting the product toward the camera.

Request

Both reference types go in input_references with their own type.

import os, time, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
    "model": "gemini-omni-flash-1.1",
    "prompt": "The person in <VIDEO_REF_0> lifts the bottle from <IMAGE_REF_0> toward the camera and smiles",
    "duration": 6,
    "resolution": "720p",
    "aspect_ratio": "9:16",
    "input_references": [
        {"type": "image_url", "image_url": {"url": "https://example.com/bottle.png"}},
        {"type": "video_url", "video_url": {"url": "https://example.com/person-3s.mp4"}},
    ],
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
    time.sleep(30)
    s = requests.get(job["polling_url"], headers=H).json()
    if s["status"] in ("completed", "failed", "cancelled"):
        break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume