Which AI video models take reference images in Sume's Videos panel?
Auto, Kling 3.0, Wan 3.0 and MiniMax H3 show reference slots in Sume's panel; H3 Max and Grok do not. Plus the API limits for image, video and audio references.

In Sume's Agents Videos panel, Auto, Kling 3.0, Wan 3.0 and MiniMax H3 show reference slots; MiniMax H3 Max and Grok Imagine do not. If your clip needs a recurring character, product or style, pick from the first four, or use the API where the catalog lists what each model accepts.
The panel and the API are not identical, and the difference trips people up. This post separates the two so you can tell which one is limiting you.
What does the panel show per model?
Each panel entry carries a flag for reference slots, and the panel hides the tray when it is off. The flag is separate from the end-frame flag, which is why H3 Max can show an end frame but no references. The code comment explains the H3 Max gap: the provider route used there has no reference-to-video endpoint, so the panel offers no Edit mode for it either.
| Model | Reference slots | Edit mode | End frame |
|---|---|---|---|
| Auto | Yes | Yes | Yes |
| Kling 3.0 | Yes | Yes | No |
| Wan 3.0 | Yes | Yes | Yes |
| MiniMax H3 | Yes | Yes | Yes |
| MiniMax H3 Max | No | No | Yes |
| Grok Imagine | No | No | No |
What does the API accept?
The video generation docs separate two modes. frame_images sets first or last frames for image-to-video, while input_references provides style or content guidance for reference-to-video. If you send both, frame_images wins and the request is image-to-video.
Which reference types a model takes is listed in supported_input_references. The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. So through the API, H3 Max does report image, video and audio references even though the panel shows no slots for it.
How many references can I send?
Counts differ by model, so read them from the catalog rather than from a general post. For the legacy Auto URL, the Video 1.0 docs give 1 to 9 reference_image_urls, 1 to 3 reference_video_urls and 1 to 3 reference_audio_urls, with audio requiring at least one image or video reference. Those numbers belong to that legacy URL, not to every pinned model.
On the vendor side, Kling's 4.0 versus 3.0 page says Kling 3.0 allows up to 7 reference assets without video and up to 4 with video, while Kling 4.0 allows up to 15 combined reference assets. Sume lists Kling 3.0, so do not plan around the 4.0 number here.
- Use public HTTPS URLs for every reference.
- Do not mix a frame image and a reference on the same request and expect both; frame images win.
- Audio references need at least one image or video reference on the legacy Auto URL.
Which should I pick for a recurring character?
For the panel, Wan 3.0 is the broadest: references, an end frame and up to 30 seconds. For the API with audio references, Seedance 2.x, Wan 3.0, H3 and H3 Max all list them. If the only thing you need is to keep a person consistent between shots, the existing post on image-to-video versus reference-to-video explains when a first frame is enough.
What should I check before a batch?
Reference images must be public HTTPS URLs in a supported format, and the docs tell you to check this first when a generation fails.
Submit one short clip first, open the finished job, and compare what you asked for with what came back. Use an Idempotency-Key on each attempt, because a replay with the same key returns the original job instead of creating and billing a second one. Only then queue the rest.
Sume reserves provider list times 1.25 when a job is submitted, and the poll response's usage.cost is the billable amount. Treat that field, not a panel estimate, as the number to budget with.
What does Sume not do?
Sume does not let you mix first-frame images and references on the same request, and it does not promise the same reference counts across models. If a reference is rejected, check the model's entry in GET /v1/videos/models before retrying with the same payload.
Sources
Related posts
More in Media tools
- Why did my video settings change when I switched model in Sume?
Sume's Videos panel clamps resolution, aspect ratio, duration and audio to the new model: 720p first, first listed length, first listed ratio. Worked examples.
- Which AI video models take an end frame in Sume's video picker?
Wan 3.0, MiniMax H3, H3 Max and Auto take an end frame in Sume's Videos panel; Kling 3.0 and Grok Imagine do not. How to check the API list too.
- YouTube caption file for a Short: with timing or without timing?
YouTube's Upload file option asks for With timing or Without timing. Which to pick for a Short, and how Sume's transcript segments and burned-in captions fit.
- YouTube Shorts max 1080p: downscale a 4K vertical clip with Sume
YouTube's Shorts help says uploads have a maximum resolution of 1080p. Downscale a 2160x3840 clip to 1080x1920 with video-trim's exact-mode output conform.
Written by Sume