Person holding your product: reference clip caps on Omni, Wan and H3
To put a person with your product in one clip, send a photo plus a short reference clip. Omni takes 3 clips of 3 s, Wan 5 clips and 15 s, H3 3 clips and 15 s.

To show a person holding your product, send the product photo and a short clip of the person as references. gemini-omni-flash-1.1 takes up to 3 reference clips of 3 seconds or less each; wan-3.0 takes up to 5 clips with 15 seconds in total; minimax-h3 and minimax-h3-max take up to 3 clips with 15 seconds in total. kling-3 and Grok take no references at all.
If you already have the person on film and want to replace them, that is a different tool: h3-max-recast swaps the people in one source video for 1 to 4 photos. Limits here are from the Video Router guide, read 2026-10-05.
Clip caps by model
Wan and H3 also list a minimum clip length and a frame-rate floor; check them against your footage before you upload.
| Model id | Reference clips | Per-clip or total limit | Notes |
|---|---|---|---|
| gemini-omni-flash-1.1 | up to 3 | each 3 s or less | 10 images, no audio references |
| wan-3.0 | up to 5 | 15 s total, 16 fps or higher | up to 10 images and 5 audio clips |
| minimax-h3 | up to 3 | 2 to 15 s each, 15 s combined | images, clips and audio together at most 12 |
| minimax-h3-max | up to 3 | 2 to 15 s each, 15 s combined | images, clips and audio together at most 12 |
Writing the prompt
On Omni, refer to the media as <IMAGE_REF_0> and <VIDEO_REF_0>, numbered from 0 in list order, and write who does what: the person in the clip picks up the item in the photo. The references are sent before the prompt, so the order of the list is the numbering.
Keep the reference clip to the part that shows the person clearly, a few seconds at most on Omni. A 3-second clip with one face and steady light is a better reference than a 15-second crowd scene.
- One person per reference clip.
- One product photo with the label facing the camera.
- Keep the prompt to one action, such as lifting the product toward the camera.
Request
Both reference types go in input_references with their own type.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
"model": "gemini-omni-flash-1.1",
"prompt": "The person in <VIDEO_REF_0> lifts the bottle from <IMAGE_REF_0> toward the camera and smiles",
"duration": 6,
"resolution": "720p",
"aspect_ratio": "9:16",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/bottle.png"}},
{"type": "video_url", "video_url": {"url": "https://example.com/person-3s.mp4"}},
],
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=H).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))Sources
Related posts
More in Use cases
- Pet adoption listing clip: MiniMax H3 at 768p for 6 seconds
A shelter dog photo becomes a 6-second clip with minimax-h3 on Sume for 45 cents at 768p; stereo audio is always on, so there is no mute switch.
- Pet product demo clip on Seedance 2.5: treat, toy and leash prompts
Three prompts for pet product demo clips on Seedance 2.5, one each for a treat, a toy and a leash, with 10 to 15 second prices at 480p, 720p and 1080p on Sume.
- Photographer portfolio reel: four portraits, four Seedance Fast clips
Four portraits, four 4-second 720p clips with seedance-2-fast on Sume cost $4.84; join them on a timeline for one reel.
- How Pinterest applies the AI label to a generated image Pin
Pinterest labels an image Pin when the owner says it is AI-made or its detection systems think so. Appeal steps are in a separate Help Center article.
Written by Sume