frame_images needs frame_type: first_frame or last_frame on Sume

Each frame_images entry on POST /v1/videos needs a frame_type. What the two values mean, how supported_frame_images limits them, and a pre-check.

5 min readSume
All posts

Every object you put in frame_images on POST /v1/videos has to say which frame it is. Sume's video docs state it plainly: each entry must include a frame_type of first_frame or last_frame. An entry with an image URL and no frame_type is an incomplete image-to-video request, and the documented class for a malformed body on Sume is 400 invalid_request, so fix the entry rather than retrying it.

This page is a short checklist for the people who hit that wall: you have a still, a model id, and a prompt, and the API wants one more field. Nothing here is a new feature; it is the part of the request shape that is easiest to miss when you port a client from a vendor that infers the frame position from list order.

What a valid frame_images entry looks like

The entry is an object with a type of image_url, an image_url object holding the public HTTPS url, and the frame_type. The docs example sends one first frame with resolution set to 1080p. Sume follows the OpenRouter Video Generation API field for field here, so a client written for that guide works after you change the base URL and the key.

Reference images are a different field. input_references carries style or content references for reference-to-video, and the model treats them as visual guidance rather than exact frames. If a request carries both fields, frame_images takes precedence and the request is treated as image-to-video, so a stray frame_images entry quietly turns a reference request into a first-frame request.

{
  "model": "seedance-2",
  "prompt": "A character walking through a forest",
  "frame_images": [
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/first-frame.png" },
      "frame_type": "first_frame"
    }
  ],
  "resolution": "1080p"
}

Which frame types does a model accept

The frame_type values you may send are model-specific. GET /v1/videos/models returns supported_frame_images per model. The docs' sample row for seedance-2 lists ["first_frame", "last_frame"], so a first-and-last-frame request is allowed there. Do not assume the same for every id. Read the field for the model you pin and send only the values it lists.

The table below is the short map between the three request ideas, read from the Sume docs on 2026-10-03.

How the two image fields differ on POST /v1/videos (read 2026-10-03)
FieldWhat it triggersEntry needsWins if both are sent
frame_imagesImage-to-video from exact framesframe_type of first_frame or last_frameYes
input_referencesReference-to-video guidanceNo frame_type; model uses it as guidanceNo
NeitherText-to-videoPrompt onlyNot applicable

A pre-check you can run before paying

Fetch the model row once, check that the frame_type you plan to send is in supported_frame_images, and fail locally if it is not. Sume reserves the cost on submit at provider list times 1.25, so a request that is going to be rejected for shape is cheaper to catch in your own code than to retry in a loop.

Keep the image URLs public and reachable over HTTPS. The troubleshooting section of the video docs says reference images must be accessible over public HTTPS and in a supported format; a private or expiring link is the next most common reason an image-to-video job fails after the request itself is accepted.

Limits of this note

The docs name the field and the rule. They do not print the exact error message text for a missing frame_type, so this post does not quote one. If you need the precise wording, send the malformed request once against a development key and read the error.details in the response.

It also does not cover which models accept a last frame today. That list changes with the catalog; the models endpoint is the source of truth.

Common mistakes with frame_images

The first mistake is sending two entries with the same frame_type. A request has one first frame and at most one last frame, so a list of three keyframes for a longer story is the wrong shape; run one job per segment and use the final frame of one clip as the first frame of the next.

The second mistake is mixing the fields by accident. A template that always includes an empty frame_images array next to real input_references is a request you may not want treated as image-to-video, so omit the field entirely when you have no frames.

The third is a missing resolution. The docs example sets resolution explicitly. A model has its own default, so say what you want and check it against supported_resolutions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume