Six product photos as references: which Sume video models accept them

Six reference photos fit Gemini Omni Flash, Wan 3.0, MiniMax H3 and H3 Max on Sume. Kling 3 and Grok Imagine Video take no references.

5 min readSume
All posts

Six reference photos of one product fit gemini-omni-flash-1.1 (up to 10 images), wan-3.0 (up to 10), minimax-h3 and minimax-h3-max (up to 9) on Sume, and the Seedance models accept image references as well. kling-3 and grok-imagine-video-1.5 take no reference media, so a six-photo brief cannot go to either of them.

Send the photos as input_references on POST /v1/videos; they condition the whole clip as visual guidance rather than as exact frames. The limits below come from the Sume catalog and the Video Router guide, read 2026-10-05.

Reference caps by model

Caps are what the catalog constraints list. A model not shown here either takes no references or has no cap stated in the catalog.

Reference inputs per Sume video model (read 2026-10-05)
Model idImage referencesVideo referencesAudio references
gemini-omni-flash-1.1up to 10up to 3, each 3 s or lessnot accepted
wan-3.0up to 10up to 5, 15 s totalup to 5, 15 s total
minimax-h3up to 9up to 3, 15 s totalup to 3, 15 s total
minimax-h3-maxup to 9up to 3, 15 s totalup to 3, 15 s total
kling-3nonenonenone
grok-imagine-video-1.5nonenonenone

Using six photos well

On Omni, refer to each photo in the prompt as <IMAGE_REF_0> to <IMAGE_REF_5>, numbered from 0 in list order. That lets the prompt say which photo is the front, the pack back or the label detail.

On the H3 rows, a reference-to-video job bills output seconds at list times 1.25 and the catalog notes that fal's reference-token overage is not reserved on H3 Max; the H3 catalog row also notes extra reference images past five carry a list add-on. Check usage.cost on the first job before you scale a batch.

  • Use clean photos with the product alone on a plain background.
  • Keep one product per reference set so the model does not blend two items.
  • Draft at the cheapest resolution the row offers, then render the approved prompt again at the final one.

Request

References go in input_references; if you also send frame_images, frame_images wins and Sume treats the job as image-to-video.

import os, time, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
    "model": "gemini-omni-flash-1.1",
    "prompt": "A hand lifts <IMAGE_REF_0> from a desk, turns it to show <IMAGE_REF_1>, natural window light",
    "duration": 6,
    "resolution": "720p",
    "aspect_ratio": "9:16",
    "input_references": [
        {"type": "image_url", "image_url": {"url": "https://example.com/front.png"}},
        {"type": "image_url", "image_url": {"url": "https://example.com/label.png"}},
    ],
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
    time.sleep(30)
    s = requests.get(job["polling_url"], headers=H).json()
    if s["status"] in ("completed", "failed", "cancelled"):
        break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume