frame_images and input_references together: frame_images wins

Send both and Sume runs image-to-video, not reference-to-video. How to pick one field per request and check the model's catalog entry first.

4 min readSume
All posts

If you send both frame_images and input_references in one POST /v1/videos body, frame_images controls the mode and Sume processes the request as image-to-video. Your references are not used as style guidance in that request. Choose one field per request, and check the model's catalog entry first.

This is documented behavior on the video generation page, but it is easy to miss when a form builder always submits both arrays.

The two fields

Descriptions are from the video generation page; the third column is what the catalog advertises per model.

frame_images vs input_references (read 2026-10-07)
FieldMode it startsCatalog field that says if a model accepts it
frame_images (each with frame_type: first_frame or last_frame)Image-to-videosupported_frame_images
input_references (image, video or audio URLs)Reference-to-video (visual guidance, not exact frames)supported_input_references

Checking a model before you send

The catalog entry for each model lists what it supports, for example supported_frame_images: ["first_frame", "last_frame"] and supported_input_references: ["image_url", "video_url", "audio_url"] on Seedance 2.0. A model only accepts a reference type its catalog entry lists, so do not send the others. Gemini Omni Flash 1.1 takes image and video references and no audio.

Media must be reachable. If Sume cannot fetch or mirror an input it returns an error such as image_not_fetchable or input_media_unreachable; the fix is a public HTTPS URL, then a retry.

Pick the field from the intent

Use frame_images when the image is the first or last frame of the clip: a product shot that must start the video. Use input_references when the image tells the model how things should look but does not need to appear as a frame: a character sheet or a palette.

A UI should make these exclusive. Offer a "start from this image" control and a "match this look" control, and clear one when the user fills the other.

{
  "model": "seedance-2",
  "prompt": "The mug turns slowly on a wooden table",
  "duration": 6,
  "frame_images": [
    {"type": "image_url", "image_url": {"url": "https://example.com/mug.jpg"}, "frame_type": "first_frame"}
  ]
}

Check the shape

The sample follows the image-to-video example shape in the docs; check the video generation page for the exact object fields before copying, because the entry format for each array is documented there and not repeated here.

Test both paths

Write two request tests per model you support: one with a first frame and no references, one with references and no frame. Assert on the model's catalog fields before sending, so a model that lacks support fails in your test and not in a paid call.

If you migrate from a vendor whose API takes a single image field, map it to frame_images with frame_type: first_frame and leave input_references empty. Only add references on models whose catalog entry lists them.

If your users can upload both kinds of image, store which one each upload is for. A single "images" array that you later split by guesswork is how a reference quietly becomes a frame.

  • Read supported_frame_images and supported_input_references from the live catalog.
  • Send public HTTPS URLs only.
  • Show users which mode their request will use.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume