Which AI video models accept an input video on Sume?
Seedance, Wan 3.0, H3, H3 Max and Gemini Omni Flash take video references; Recast, Genjutsu and Omni edit need a source video. Kling and Grok take none.

There are two different things a video can be to a Sume model. A reference video is extra context that guides a new clip, such as motion or style. A source video is the clip being transformed, and the output keeps its length. They are not interchangeable, and a row that supports one often rejects the other with a 400.
The split is advertised in the catalog. GET /v1/videos/models returns supported_input_references per row, and the video-router capability flags say whether the row can run video-to-video.
Which rows take which?
The table summarizes the catalog. Kling Video v3 Pro and Grok Imagine Video 1.5 accept no input video at all; Grok is image-to-video only.
| Row | What it accepts | Input video |
|---|---|---|
seedance-2.5 and Seedance 2.0 family | video_url reference | Yes (reference) |
wan-3.0 | reference_video_urls, up to 5, 15 s total | Yes (reference) |
minimax-h3, minimax-h3-max | video_url reference | Yes (reference) |
gemini-omni-flash-1.1 | Up to 3 reference clips, each 3 s or less; or video_url edit | Yes (reference, edit) |
h3-max-recast | One source video plus 1 to 4 photos | Required |
higgsfield-genjutsu | One source video plus 1 to 8 photos | Required |
kling-3, grok-imagine-video-1.5 | None | No |
What are the limits on references?
The limits differ per row, so check them before you upload. Wan 3.0 allows up to five reference videos of at most 15 seconds in total, at 16 frames per second or more. Gemini Omni Flash 1.1 allows up to three reference clips, each 3 seconds or shorter. The other rows do not publish a count in the constraints Sume exposes, so test with the catalog before relying on more than one.
On /v1/videos the references go in input_references as video_url entries. On the legacy Video Router wire, source clips use video_url, and Recast and Genjutsu photos use reference_image_urls.
How do I pick?
Want to keep a source clip's timing and change who is in it: Recast or Genjutsu. Want to edit a clip with a prompt: Omni's edit path on the Video Router. Want a new clip that borrows motion from an example: a row with reference videos. For a text-only workflow, Kling is the simpler choice, since it takes no references to get wrong.
Sources
Related posts
More in Models
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which AI video models take a reference audio file on Sume?
Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max accept a reference audio clip on Sume; Gemini Omni Flash, Kling, Grok, Recast and Genjutsu do not.
- Which Sume image models accept quality? Only five do
Only five Sume image catalog rows list a quality field: GPT Image 2, 2.5 and Sunburst, Ideogram V3 and 4.5. The rest return 400 unsupported_parameter.
- Which Sume image models can't edit? Five text-to-image-only rows
Five Sume image rows take text prompts only and reject image_urls: Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max. Prices and what to use instead.
Written by Sume