Sent frame_images and input_references together? frame_images win

If a /v1/videos request has both fields, Sume treats it as image-to-video and the reference images do not condition the clip. How to pick one, with JSON bodies.

5 min readSume
All posts

If a /v1/videos request on Sume carries both frame_images and input_references, frame_images wins and the call is image-to-video. The references are not used as guidance. Send frame_images to pin an opening or closing frame, and input_references to keep a subject or style across a free-form clip.

How the API picks a mode

The caller never names the mode. The contract in docs/api/videos.md infers it: frame_images present gives image-to-video; input_references present gives reference-to-video; neither gives text-to-video; both gives image-to-video. The design came from bug #2044, where an ambiguous single image field could pin a frame when the caller meant a reference.

Pin a frame

Each entry in frame_images needs a frame_type of first_frame or last_frame. A model accepts the type only if supported_frame_images in the catalog lists it. Seedance 2.5 lists both, so you can give an opening and a closing frame.

{
  "model": "seedance-2.5",
  "prompt": "Slow push in on the sneaker, soft studio light",
  "duration": 6,
  "resolution": "720p",
  "frame_images": [
    { "type": "image_url", "image_url": { "url": "https://example.com/shoe-start.png" }, "frame_type": "first_frame" }
  ]
}

Keep a subject

A single image in input_references conditions the whole clip rather than becoming the first frame. That is the fix for #2044. Reference limits depend on the model: Wan 3.0 takes up to 10 images, 5 videos and 5 audio files, and Gemini Omni Flash 1.1 takes up to 10 images and 3 clips of at most 3 s with no audio. The table in our reference-limits note has the rest.

{
  "model": "wan-3.0",
  "prompt": "The mug from the reference photo steams on a wooden desk",
  "duration": 5,
  "input_references": [
    { "type": "image_url", "image_url": { "url": "https://example.com/mug.png" } }
  ]
}

Both at once: what to do

If you want a subject locked and a pinned frame, put the subject photo in frame_images as the first frame, and describe the rest in the prompt. A model cannot take both modes at once through this surface, and a silent downgrade of a reference set is the sort of mistake that costs a full render. Look at the poll response model and your request body before you submit a batch.

On the legacy Video Router the field names differ (image_url, end_image_url, reference_image_urls), but the same rule holds: a single reference image with no frame is reference-to-video.

Catching the mistake early

Add a check in your client: if both arrays are non-empty, fail locally. The server will not tell you, because the request is valid. In the poll response, model and status do not show the mode, so the only sign is a clip that ignores your references.

The catalog fields help. supported_frame_images lists the frame_type values, and supported_input_references lists the allowed reference types. Check both before you build a request for a new model.

Related posts

More in Developers

All Developers posts

Written by Sume