First frame plus style sheet in one video call: frame_images wins
On Sume /v1/videos, frame_images plus input_references runs as image-to-video, so plan on your style sheet being set aside. Split the call and pin the opening.

If you send both frame_images and input_references in one POST /v1/videos request, Sume treats it as image-to-video: frame_images takes precedence, so the images you meant as a style sheet are not used as references. For an episode that needs an exact opening and a consistent look, make two decisions instead of one call: pin the opening with a first frame, and carry the look in the prompt or in the still itself.
This comes up as soon as a series gets an image workflow. xAI's image guide describes grok-imagine-image-2.0 generating up to 10 images per request and editing from up to 5 source images (read 2026-10-03), which makes it easy to end up with a character sheet, a location still and an approved opening all in hand. The temptation is to hand all of them to the video model at once.
Two modes, not one
The video docs describe two ways to provide images, and each triggers a different generation mode. frame_images is for image-to-video: each entry needs a frame_type of first_frame or last_frame. input_references is for reference-to-video: the model uses the images as visual guidance rather than exact frames.
The rule when both are present is plain: frame_images takes precedence and the request is treated as image-to-video. Nothing in that rule says the references are merged. Plan on them being ignored, and check the result against your expectations rather than assuming a blend.
| Field | Mode | What the image is | Capability flag |
|---|---|---|---|
| frame_images | Image-to-video | An exact first or last frame | supported_frame_images |
| input_references | Reference-to-video | Style or content guidance | supported_input_references |
| Both together | Image-to-video | References do not apply | frame_images wins |
Check what the model accepts
Both lists are published per model. GET /v1/videos/models returns supported_frame_images (for example first_frame and last_frame) and supported_input_references (image, video or audio types). A model that does not list a type will not accept it, so read the entry for the model you plan to use rather than copying a request that worked on another one.
The video docs also say audio and video references are honoured by some model families and not others, and that some accept video references but not audio. Treat that as a reason to look at the live list, not to memorise it.
import os, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
wanted = os.environ.get("VIDEO_MODEL", "seedance-2.5")
models = requests.get(f"{API}/v1/videos/models", headers=H).json()
items = models.get("data", models.get("models", []))
for m in items:
if m.get("id") == wanted:
print("frame images:", m.get("supported_frame_images"))
print("input refs: ", m.get("supported_input_references"))
break
else:
print("model id not in the list; check GET /v1/catalog")What to do for an episode
Decide which job each image has. The approved opening still is a first frame. A character sheet or palette is a reference. If both matter and the model is image-to-video, bake the look into the opening still itself, which is cheap to redo, and describe the motion in the prompt. A still is a fraction of a clip's cost, so iterate there.
If the look matters more than the exact opening, drop the first frame and use references alone, then pick the best frame from the result with video frames and use it as the next episode's start. That keeps the series coherent without fighting the precedence rule.
Test once per model with a throwaway request before you queue a season. Send the same prompt with only a first frame, and again with only references, and compare. It costs two short clips and tells you whether the model keeps the face or the framing better. Use an idempotency key in each so a retry does not bill twice.
Write the choice into your episode template so nobody re-adds the second field later. A request builder that sets one of the two and refuses to set both is a few lines, and it removes a class of silent surprises.
What Sume does and does not do
Sume documents the precedence and publishes per-model support, and it runs the request you send. It does not merge a frame and a style sheet into one conditioned clip, and the docs do not say it flags references that were set aside; that is why the check belongs in your code. Vendor image limits such as the five edit sources above describe xAI's own API, not Sume's.
Sources
Related posts
More in Models
- Shortest AI video clip via API: 2 seconds on Wan 3.0 only
Sume's video catalog by minimum duration: Wan 3.0 takes 2 seconds, Omni 3, Seedance and Kling 4, MiniMax H3 5. Plus how to trim below a floor.
- Shortest AI video clip by model: 2, 3, 4 or 5 second floors on Sume
Wan 3.0 starts at 2 seconds; Gemini Omni Flash at 3; Seedance and Kling at 4; MiniMax H3 at 5. The minimum duration for each Sume video model.
- Shortest AI video clip each Sume model accepts, and its price
Wan 3.0 starts at 2 seconds ($0.125 at 480p), Omni Flash at 3, Seedance 2.5 and Genjutsu at 4, H3 at 5. Minimum request per model, priced.
- Silent AI video: generate_audio false or drop the audio after
Seedance, Kling and Wan take generate_audio false; Omni and MiniMax H3 always make sound. How to get a silent clip on Sume, and what it costs.
Written by Sume