Video catalog input matrix: which model takes frames or references
Flatten GET /v1/videos/models into one table of first and last frame, image, video and audio reference support plus audio, in a short Python script.

To see which Sume video model accepts a first frame, a last frame, or image, video and audio references, read three fields from GET /v1/videos/models: supported_frame_images, supported_input_references and generate_audio. A short Python script turns the response into one matrix, so you stop guessing from launch posts.
Launch posts list headline features such as reference counts, but your key can only call what the catalog lists. Sume's video generation docs (read 2026-10-02) say limits are not uniform across models, so a matrix built from the live response is the safe starting point before you write request code.
What do the three fields mean?
Each model entry carries the descriptors below. They map one-to-one onto request fields, so the matrix tells you what a request may contain.
One rule matters when you combine inputs: if you send both frame_images and input_references, frame_images takes precedence and the request is treated as image-to-video. A matrix row that lists both capabilities does not mean the two can work together in one call.
| Catalog field | Example values in the docs | Request field it gates |
|---|---|---|
supported_frame_images | first_frame, last_frame | frame_images[].frame_type |
supported_input_references | image_url, video_url, audio_url | input_references[] entries by type |
generate_audio | true or false | generate_audio |
supported_durations | whole seconds, such as 4 to 15 | duration |
How do I build the matrix in Python?
The script below requests the catalog once and prints one line per model. Missing fields are treated as empty lists, because a model that cannot take a type should not appear to take it.
import os
import requests
resp = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
resp.raise_for_status()
print(f"{'model':26} {'frames':24} {'refs':28} audio")
for m in resp.json()["data"]:
frames = ",".join(m.get("supported_frame_images") or []) or "-"
refs = ",".join(m.get("supported_input_references") or []) or "-"
audio = m.get("generate_audio")
print(f"{m['id']:26} {frames:24} {refs:28} {audio}")What do the docs say about specific models?
The docs name several limits you can check against your output. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. Gemini Omni Flash 1.1 also exposes its video edit mode through a video_url field on the Video Router.
If your live output disagrees with the docs, trust the response and tell your team: some models are listed only when their provider is configured, so your catalog can differ from the docs.
What should I do with the matrix?
Save it as a CSV next to your prompts. When a vendor announces a feature, such as a new reference type, look for the matching descriptor before you plan work around it. If the descriptor is absent, the feature is not callable through Sume yet, whatever the launch page says.
Read the catalog again after each launch week and diff the saved file; the related post on checking the catalog after a launch covers that habit.
Sources
Related posts
More in Developers
- low_confidence_long_video: why video_frames warns past 90 seconds
Sume's video_frames returns the low_confidence_long_video warning when the source runs over 90 s. The job still succeeds; the hard cap is 300 s. What to do.
- Check has_audio first: video_inspect frames false before STT or detach
A free probe-only video_inspect tells you probe.has_audio before you reserve STT or run audio detach, so silent clips never hit the no-audio errors.
- video_inspect silence_split_seconds: sentence segments for captions
How silence_split_seconds (0.2 to 3) shapes Sume video-inspect sentence segments, the 0.5 s default in the repo, and turning segments into caption cues.
- Video over 30 minutes: Sume inspect caps at 1800 s
Sume video inspect and video trim reject sources over 1800 seconds. For a longer recording, split it before import, or use audio detach and STT.
Written by Sume