Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.

Pick the Sume video model by the inputs you already have. Text only: any text-to-video row except grok-imagine-video-1.5. One photo as the first frame: any image-to-video row. A first and a last frame: seedance-2.5, seedance-2, kling-3 or gemini-omni-flash-1.1, not Grok. Several photos or clips as references: the Seedance, Wan, MiniMax H3 and Omni rows, not Kling. A voice or sound reference: Seedance 2.x, Wan 3.0 or MiniMax H3, not Omni. An existing video to change: gemini-omni-flash-1.1, h3-max-recast or higgsfield-genjutsu.
The rules come from Sume's Video Router docs, Video generation docs and OpenAPI request schema, read on 2026-10-03. The catalog changes, so the last section shows how to confirm any row with one call.
Which model matches each input?
Start with the row for the strongest input you have and work down. Frames and references are separate modes. On /v1/videos, if both frame_images and input_references are sent, the frames win and the request is treated as image-to-video, and on some models a first frame plus references is rejected outright.
| You have | Send | Models that accept it |
|---|---|---|
| A prompt only | prompt | Seedance 2.5, 2, Fast and Mini; Kling 3; Wan 3.0; MiniMax H3 and H3 Max; Omni |
| One first-frame photo | image_url | Every text-capable row above, and grok-imagine-video-1.5, which requires it |
| First and last frame | image_url + end_image_url | Seedance 2.5, Seedance 2, Kling 3, Omni; not Grok |
| Several photos as references | reference_image_urls | Seedance family, Wan 3.0, MiniMax H3, H3 Max, Omni (up to 10); not Kling, not Grok |
| Reference clips | reference_video_urls | Seedance 2.x, Wan 3.0, MiniMax H3, H3 Max, Omni (3 clips, 3 s each) |
| A voice or sound sample | reference_audio_urls | Seedance 2.x, Wan 3.0, MiniMax H3, H3 Max; not Omni |
| A clip to edit with a prompt | video_url | gemini-omni-flash-1.1 only |
| A clip plus 1-4 person photos | video_url + reference_image_urls | h3-max-recast |
| A clip plus 1-8 reference images | video_url + reference_image_urls | higgsfield-genjutsu, listed only when its provider is configured |
What does a first-and-last-frame request look like?
This is the two-photo case from the table, pinned to kling-3, which accepts 720p and 1080p and 4 to 15 seconds. The last frame cannot be sent alone: end_image_url requires image_url. Use public HTTPS image URLs.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: kling-frames-001" \
-d '{
"model": "kling-3",
"prompt": "Slow dolly from the empty shelf to the stocked shelf",
"image_url": "https://example.com/inputs/shelf-empty.png",
"end_image_url": "https://example.com/inputs/shelf-full.png",
"resolution": "1080p",
"duration": 5,
"aspect_ratio": "16:9",
"mode": "async"
}'What if I have no strong preference?
Use sume/auto on POST /v1/videos. Auto creates 3 to 10 second clips, defaults to 720p and 8 seconds, and works at 16:9 or 9:16. It does not tell you which model served the clip, so it fits cases where the look matters less than getting a clip. When you need to know the model, or a ratio or length Auto lacks, pin one from the table.
sume/auto: no square, no 21:9, which model to pin lists what Auto cannot do, and image-to-video vs reference-to-video explains why a first frame and a reference set behave differently.
Which input combinations does Sume reject?
The schema is strict about a few mixes, and the errors are fast, so you find out at submit rather than after a render.
end_image_urlneedsimage_url; a last frame alone is rejected.grok-imagine-video-1.5rejects references, an end frame andgenerate_audio, and requiresimage_url.kling-3rejectsreference_*_urls; use a Seedance, Wan, MiniMax or Omni row for references.- On Omni,
video_urlcannot be combined withimage_url,end_image_urlorreference_*_urls, andgenerate_audio: falseis rejected. - Audio alone is not a valid reference set on MiniMax H3 or H3 Max; send at least one image or video with it.
h3-max-recastandhiggsfield-genjutsuneedvideo_urltogether with reference images, and the source length sets the output length.
How do I confirm a row before I build on it?
Ask the catalog. GET /v1/video-router/models returns capabilities per model, and GET /v1/videos/models returns supported_resolutions, supported_aspect_ratios, supported_durations, supported_frame_images and supported_input_references. Sume's own docs say to read those fields rather than assume one envelope, and they are the first thing to diff when a new model appears; check the Sume catalog with one call shows the request.
Treat a table like the one above as a map, not a contract. A model that gains a mode appears in the catalog before it appears in a blog post.
Sources
Related posts
More in Models
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
- Music generation API: the Sume Music Router with Lyria 3.5
Sume's Music Router turns a text prompt into a track via POST /v1/music-router/generate. sume/music-auto picks the engine, Lyria 3.5 today.
Written by Sume