frame_images and input_references together: frame_images wins
Send both and Sume runs image-to-video, not reference-to-video. How to pick one field per request and check the model's catalog entry first.

If you send both frame_images and input_references in one POST /v1/videos body, frame_images controls the mode and Sume processes the request as image-to-video. Your references are not used as style guidance in that request. Choose one field per request, and check the model's catalog entry first.
This is documented behavior on the video generation page, but it is easy to miss when a form builder always submits both arrays.
The two fields
Descriptions are from the video generation page; the third column is what the catalog advertises per model.
| Field | Mode it starts | Catalog field that says if a model accepts it |
|---|---|---|
frame_images (each with frame_type: first_frame or last_frame) | Image-to-video | supported_frame_images |
input_references (image, video or audio URLs) | Reference-to-video (visual guidance, not exact frames) | supported_input_references |
Checking a model before you send
The catalog entry for each model lists what it supports, for example supported_frame_images: ["first_frame", "last_frame"] and supported_input_references: ["image_url", "video_url", "audio_url"] on Seedance 2.0. A model only accepts a reference type its catalog entry lists, so do not send the others. Gemini Omni Flash 1.1 takes image and video references and no audio.
Media must be reachable. If Sume cannot fetch or mirror an input it returns an error such as image_not_fetchable or input_media_unreachable; the fix is a public HTTPS URL, then a retry.
Pick the field from the intent
Use frame_images when the image is the first or last frame of the clip: a product shot that must start the video. Use input_references when the image tells the model how things should look but does not need to appear as a frame: a character sheet or a palette.
A UI should make these exclusive. Offer a "start from this image" control and a "match this look" control, and clear one when the user fills the other.
{
"model": "seedance-2",
"prompt": "The mug turns slowly on a wooden table",
"duration": 6,
"frame_images": [
{"type": "image_url", "image_url": {"url": "https://example.com/mug.jpg"}, "frame_type": "first_frame"}
]
}
Check the shape
The sample follows the image-to-video example shape in the docs; check the video generation page for the exact object fields before copying, because the entry format for each array is documented there and not repeated here.
Test both paths
Write two request tests per model you support: one with a first frame and no references, one with references and no frame. Assert on the model's catalog fields before sending, so a model that lacks support fails in your test and not in a paid call.
If you migrate from a vendor whose API takes a single image field, map it to frame_images with frame_type: first_frame and leave input_references empty. Only add references on models whose catalog entry lists them.
If your users can upload both kinds of image, store which one each upload is for. A single "images" array that you later split by guesswork is how a reference quietly becomes a frame.
- Read
supported_frame_imagesandsupported_input_referencesfrom the live catalog. - Send public HTTPS URLs only.
- Show users which mode their request will use.
Sources
Related posts
More in Developers
- Gemini 3.8 Flash TTS takes 8,192 input tokens: splitting a long script
Gemini 3.8 Flash TTS lists 8,192 input tokens and 16,384 output tokens. Sume TTS takes 20,000 characters per job. Here is how to split a long script.
- Gemini TTS streams raw PCM; Sume returns a finished wav or raw file
Gemini 3.8 Flash TTS streams headerless audio/l16 at 24 kHz. Sume returns a hosted file from an async job. Wrap the PCM in WAV and match the formats.
- Omni Flash 4K, 10 seconds: a webhook-mode /v1/videos request body
One curl body for gemini-omni-flash-1.1 at 4K, 16:9, 10 seconds with audio and a callback_url, plus what the docs say the submit reserves and what it cannot.
- Gemini Omni Flash edit on Sume: a Python preflight for bad fields
A 29-line Python preflight for Sume video-router edit requests on gemini-omni-flash-1.1: catches the fields the docs say it rejects before you pay for a job.
Written by Sume