frame_images needs frame_type: first_frame or last_frame on Sume
Each frame_images entry on POST /v1/videos needs a frame_type. What the two values mean, how supported_frame_images limits them, and a pre-check.

Every object you put in frame_images on POST /v1/videos has to say which frame it is. Sume's video docs state it plainly: each entry must include a frame_type of first_frame or last_frame. An entry with an image URL and no frame_type is an incomplete image-to-video request, and the documented class for a malformed body on Sume is 400 invalid_request, so fix the entry rather than retrying it.
This page is a short checklist for the people who hit that wall: you have a still, a model id, and a prompt, and the API wants one more field. Nothing here is a new feature; it is the part of the request shape that is easiest to miss when you port a client from a vendor that infers the frame position from list order.
What a valid frame_images entry looks like
The entry is an object with a type of image_url, an image_url object holding the public HTTPS url, and the frame_type. The docs example sends one first frame with resolution set to 1080p. Sume follows the OpenRouter Video Generation API field for field here, so a client written for that guide works after you change the base URL and the key.
Reference images are a different field. input_references carries style or content references for reference-to-video, and the model treats them as visual guidance rather than exact frames. If a request carries both fields, frame_images takes precedence and the request is treated as image-to-video, so a stray frame_images entry quietly turns a reference request into a first-frame request.
{
"model": "seedance-2",
"prompt": "A character walking through a forest",
"frame_images": [
{
"type": "image_url",
"image_url": { "url": "https://example.com/first-frame.png" },
"frame_type": "first_frame"
}
],
"resolution": "1080p"
}Which frame types does a model accept
The frame_type values you may send are model-specific. GET /v1/videos/models returns supported_frame_images per model. The docs' sample row for seedance-2 lists ["first_frame", "last_frame"], so a first-and-last-frame request is allowed there. Do not assume the same for every id. Read the field for the model you pin and send only the values it lists.
The table below is the short map between the three request ideas, read from the Sume docs on 2026-10-03.
| Field | What it triggers | Entry needs | Wins if both are sent |
|---|---|---|---|
| frame_images | Image-to-video from exact frames | frame_type of first_frame or last_frame | Yes |
| input_references | Reference-to-video guidance | No frame_type; model uses it as guidance | No |
| Neither | Text-to-video | Prompt only | Not applicable |
A pre-check you can run before paying
Fetch the model row once, check that the frame_type you plan to send is in supported_frame_images, and fail locally if it is not. Sume reserves the cost on submit at provider list times 1.25, so a request that is going to be rejected for shape is cheaper to catch in your own code than to retry in a loop.
Keep the image URLs public and reachable over HTTPS. The troubleshooting section of the video docs says reference images must be accessible over public HTTPS and in a supported format; a private or expiring link is the next most common reason an image-to-video job fails after the request itself is accepted.
Limits of this note
The docs name the field and the rule. They do not print the exact error message text for a missing frame_type, so this post does not quote one. If you need the precise wording, send the malformed request once against a development key and read the error.details in the response.
It also does not cover which models accept a last frame today. That list changes with the catalog; the models endpoint is the source of truth.
Common mistakes with frame_images
The first mistake is sending two entries with the same frame_type. A request has one first frame and at most one last frame, so a list of three keyframes for a longer story is the wrong shape; run one job per segment and use the final frame of one clip as the first frame of the next.
The second mistake is mixing the fields by accident. A template that always includes an empty frame_images array next to real input_references is a request you may not want treated as image-to-video, so omit the field entirely when you have no frames.
The third is a missing resolution. The docs example sets resolution explicitly. A model has its own default, so say what you want and check it against supported_resolutions.
Sources
Related posts
More in Developers
- Gemini CLI settings.json httpUrl, timeout and includeTools for Sume
Gemini CLI's MCP entry takes httpUrl, headers, a 600000 ms timeout and includeTools. Set up Sume's hosted server with a read-only tool list and one credential.
- Gemini Live Translate transcripts as subtitles: you supply timing
Google's Live Translate can return input and output transcripts, but the page lists no word times. To burn subtitles, time each line, then send cues to Sume.
- Gemini Omni edit 400: aspect_ratio is not supported, framing is kept
Sending aspect_ratio with a video_url edit on gemini-omni-flash-1.1 returns a 400 on Sume. Why the output keeps the source framing and what to send.
- Omni edit came back 720p: how to ask for 1080p on Sume
A gemini-omni-flash-1.1 edit defaults to 720p when you leave resolution out. Set resolution on the request; here is what the edit mode accepts.
Written by Sume