Choose an AI video model by the asset you already have: Sume ids
Photo, two frames, product shots, a voice track, a driving clip or a finished video: which Sume video id fits each starting asset, from the catalog rows.

Pick the video model from the asset you already hold. A single photo goes to an image-to-video row such as wan-3.0, minimax-h3 or kling-3. A start and end frame pair works on rows with end-frame support. Several product shots need a reference row. A voice track needs a row that accepts audio references. A driving clip with a photo goes to Kling Motion Control, and a finished video to edit goes to gemini-omni-flash-1.1.
This follows Sume's catalog capabilities, which you can read live from the Video Router; it is a map of what is accepted, not a quality ranking.
The map
Capabilities come from the catalog rows checked on 2026-10-04. A row can be accepted and still be the wrong style for your brand, so test one short clip first.
| You have | Fits | Does not fit |
|---|---|---|
| One photo | wan-3.0, minimax-h3, minimax-h3-max, seedance-2.5, kling-3, gemini-omni-flash-1.1, grok-imagine-video-1.5 | - |
| Start and end frame | wan-3.0, minimax-h3, minimax-h3-max, seedance-2.5, kling-3, gemini-omni-flash-1.1 | grok-imagine-video-1.5 |
| Up to 9 or 10 product or character images | wan-3.0 (10), minimax-h3 and minimax-h3-max (9), gemini-omni-flash-1.1 (10), seedance-2.5 (read its limit from the catalog) | kling-3 |
| A voice or music track as a reference | wan-3.0, seedance-2.5, minimax-h3, minimax-h3-max | gemini-omni-flash-1.1, kling-3 |
| A video to edit by prompt | gemini-omni-flash-1.1 with video_url | every other text or image row |
| People to swap in a clip | h3-max-recast | - |
| A photo and a driving video | Kling 3.0 Motion Control route | text and image rows |
Three common choices
Most buyers fall into one of these cases, and each has a known limit on Sume.
- Ad with several product photos:
wan-3.0takes 10 images, 5 videos and 5 audio files; the MiniMax rows take 9, 3 and 3. - Short social loop from one photo:
minimax-h3takes 5 to 15 seconds at native 480p or 768p; compare per-second list prices in the catalog before choosing. - Fix one shot in a clip:
gemini-omni-flash-1.1edit, with the source framing kept and no aspect ratio to set.
Check the live row
Because rows change, ask the catalog for the flags instead of copying this table. This prints every row that accepts reference audio.
import os
import requests
r = requests.get(
"https://api.sume.com/v1/video-router/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for row in r.json()["data"]["models"]:
if row.get("capabilities", {}).get("reference_audios"):
print(row["id"])Where each row is documented
The Video Router page covers the explicit catalog and its limits, and Video generation covers the OpenRouter-compatible /v1/videos route over the same catalog. For Kling Motion Control from an agent, the hosted tool is kling-motion-control_create on the MCP tools page.
Sources
Related posts
More in Models
- Swap the gift in a Christmas clip with Gemini Omni Flash edit
Gemini Omni Flash 1.1 on Sume edits a 3 to 10 second clip by prompt, so one approved Christmas ad can become three gift variants without a reshoot.
- Does Lyria 3.5 audio carry a SynthID watermark?
Yes: Google's Gemini API music docs say Lyria audio carries a SynthID watermark. The Sume docs I read do not describe watermarking either way.
- Eleven v4 stacked tags: direct emotion without them
ElevenLabs v4 adds stackable expression tags and 10-second cloning. Sume's docs list no clone route; here is how to direct a voice via tts_create.
- eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.
Written by Sume