One reference image on Seedance 2.5 is not a first frame

On /v1/videos, input_references steer a generation; frame_images pin frames. Send one image the wrong way and the clip does not start on your picture.

5 min readSume
All posts

On POST /v1/videos, a single image in input_references makes a reference-to-video request for seedance-2.5, not an image-to-video one, so the clip does not have to begin on your picture. To make the image the opening frame, send it in frame_images with frame_type: "first_frame". If you send both fields, frame_images takes precedence and the references are not used.

The Video generation docs describe the split: frame_images specifies first or last frames for image-to-video, and input_references provides style or content references that the model uses as visual guidance rather than exact frames.

The same image, two requests

First, the image as a first frame. This is the one for a product still that must be the opening shot:

``json { "model": "seedance-2.5", "prompt": "The mug slowly rotates, steam rising", "frame_images": [ { "type": "image_url", "image_url": { "url": "https://example.com/mug.png" }, "frame_type": "first_frame" } ], "resolution": "720p", "duration": 8, "aspect_ratio": "9:16" } ``

Second, the same image as a reference. The model treats it as guidance for the look of the subject and may open on a different composition:

``json { "model": "seedance-2.5", "prompt": "A ceramic mug on a desk, slow orbit, steam rising", "input_references": [ { "type": "image_url", "image_url": { "url": "https://example.com/mug.png" } } ], "resolution": "720p", "duration": 8, "aspect_ratio": "9:16" } ``

Rules that trip people up

Each frame_images entry needs a frame_type. A last_frame without a first_frame is rejected with unsupported_capability and the message that a last frame also requires a first frame. Only types listed in a model's supported_input_references are accepted, and the Seedance 2.x family honors image, video and audio references.

On the Video Router wire the same split uses flat fields: image_url for the first frame and reference_image_urls for references. That route keeps its own envelope, so the Video generation docs recommend /v1/videos for new code.

Which field does what on Sume (read 2026-10-03)
Goal/v1/videos fieldVideo Router field
Start exactly on this imageframe_images with first_frameimage_url
End on this image (needs a first frame)frame_images with last_frameend_image_url
Match the look of this subjectinput_references image_urlreference_image_urls

Why the distinction exists

The code that builds the provider request carries a comment explaining it: a single reference image is still reference-to-video, because coercing it into image-to-video pinned frame 0 and dropped the reference conditioning. A lone input_references image therefore keeps its role as guidance. Sume does not quietly turn it into a first frame.

Limits also differ by mode. Seedance accepts up to 12 reference files combined, across images, videos and audio, while Wan 3.0 counts per type: 10 images, 5 videos and 5 audio files. Frame-based requests carry at most a first and a last frame.

A quick checklist

If your clip ignores your product shot, check whether the picture went into input_references. If you want both an exact opening frame and extra style guidance, remember that the frame fields win and the references are dropped, so choose one mode per request. For Seedance 2.5 the allowed durations are 4 to 30 seconds, so an 8-second test costs less than a 30-second one while you find out which field gives the result you wanted.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume