Six product photos as references: which Sume video models accept them
Six reference photos fit Gemini Omni Flash, Wan 3.0, MiniMax H3 and H3 Max on Sume. Kling 3 and Grok Imagine Video take no references.

Six reference photos of one product fit gemini-omni-flash-1.1 (up to 10 images), wan-3.0 (up to 10), minimax-h3 and minimax-h3-max (up to 9) on Sume, and the Seedance models accept image references as well. kling-3 and grok-imagine-video-1.5 take no reference media, so a six-photo brief cannot go to either of them.
Send the photos as input_references on POST /v1/videos; they condition the whole clip as visual guidance rather than as exact frames. The limits below come from the Sume catalog and the Video Router guide, read 2026-10-05.
Reference caps by model
Caps are what the catalog constraints list. A model not shown here either takes no references or has no cap stated in the catalog.
| Model id | Image references | Video references | Audio references |
|---|---|---|---|
| gemini-omni-flash-1.1 | up to 10 | up to 3, each 3 s or less | not accepted |
| wan-3.0 | up to 10 | up to 5, 15 s total | up to 5, 15 s total |
| minimax-h3 | up to 9 | up to 3, 15 s total | up to 3, 15 s total |
| minimax-h3-max | up to 9 | up to 3, 15 s total | up to 3, 15 s total |
| kling-3 | none | none | none |
| grok-imagine-video-1.5 | none | none | none |
Using six photos well
On Omni, refer to each photo in the prompt as <IMAGE_REF_0> to <IMAGE_REF_5>, numbered from 0 in list order. That lets the prompt say which photo is the front, the pack back or the label detail.
On the H3 rows, a reference-to-video job bills output seconds at list times 1.25 and the catalog notes that fal's reference-token overage is not reserved on H3 Max; the H3 catalog row also notes extra reference images past five carry a list add-on. Check usage.cost on the first job before you scale a batch.
- Use clean photos with the product alone on a plain background.
- Keep one product per reference set so the model does not blend two items.
- Draft at the cheapest resolution the row offers, then render the approved prompt again at the final one.
Request
References go in input_references; if you also send frame_images, frame_images wins and Sume treats the job as image-to-video.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
"model": "gemini-omni-flash-1.1",
"prompt": "A hand lifts <IMAGE_REF_0> from a desk, turns it to show <IMAGE_REF_1>, natural window light",
"duration": 6,
"resolution": "720p",
"aspect_ratio": "9:16",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/front.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/label.png"}},
],
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=H).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))Sources
Related posts
More in Media tools
- Put a slide above an AI presenter with Timeline compose: $0.02 a shot
Stack a still slide on top and a talking avatar video below in one MP4 with Sume Timeline compose. The flat price is $0.02 per job. Request and layout inside.
- Sneaker drop teaser: six clips, slideup cuts, one Timeline render
Cut six sneaker clips into one 18-second drop teaser with slideup transitions and a music bed: one Sume Timeline render, $0.10, with a free plan call first.
- Soft drop shadow under an AI product cutout with Pillow
Give a transparent product PNG from Sume a soft drop shadow: copy the alpha, blur it, offset it, tint it, then composite. Parameters for subtle and heavy looks.
- soundtrack_fade_exceeds_output: a bed fade longer than a short clip
Sume rejects soundtrack.fade_out_seconds longer than audio.duration_seconds. The fade also caps at 10 s. For a 6 s Short, use a 1 to 2 s fade.
Written by Sume