Image-to-video not starting on my photo: frame_images vs references

Your photo is a reference, not a first frame, when it goes in input_references. Use frame_images with first_frame on Sume /v1/videos to pin the opening shot.

5 min readSume
All posts

If your image-to-video clip does not open on your photo, the photo almost certainly went into the wrong field. On Sume POST /v1/videos, an image in frame_images with frame_type: "first_frame" is pinned as the opening frame. The same image in input_references is treated as visual guidance for reference-to-video, so the model is free to compose a different opening shot that only resembles it.

Sume infers the mode from which field is present and never asks you to declare it. That is convenient, but it also means a field mix-up does not return an error: you get a valid clip that just does not start where you expected. The rest of this page is a short way to tell which mode you actually requested, and how to fix the request.

Which field starts which mode

The video generation guide describes two ways to send images. The mode follows the field, so read the request you sent against this table before you blame the model.

Field to mode on /v1/videos (Sume docs, read 2026-10-07)
You sentMode the API infersWhat the image does
frame_images with first_frameImage-to-videoPinned as the opening frame
frame_images with last_frameImage-to-videoPinned as the closing frame, on models that list last_frame
input_references (image_url)Reference-to-videoStyle or content guidance, not an exact frame
Both fields in one bodyImage-to-videoframe_images wins, input_references are not the mode
NeitherText-to-videoNo image used

Fix the request

Move the photo out of input_references and into frame_images. Here is a minimal request that pins one photo as the first frame of a Seedance 2.0 clip. Replace the URL with a publicly reachable HTTPS image, because the API downloads it. If the image is behind a login, a signed link that has expired, or a private bucket, the job fails with a message about an input media URL that could not be downloaded, so test the link in a private browser window first.

Two smaller points save a retry. First, aspect ratio: the pinned frame sets the composition, so choose an aspect_ratio that the model lists and that matches your image, instead of expecting the model to crop it for you. Second, one frame is enough for a first-frame pin. Add a last_frame entry only when the model lists it and you want to control the ending as well.

The legacy Video Router surface follows the same rule with flat fields. Its Gemini Omni Flash 1.1 table says that a single reference image with no frame is reference-to-video, which is the same rule that was applied to Seedance after issue 2044. If you use /v1/video-router/generate, the field to use for a pinned opening frame is image_url, not reference_image_urls.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seedance-2",
    "prompt": "The camera slowly pushes in as steam rises from the cup",
    "resolution": "720p",
    "frame_images": [{
      "type": "image_url",
      "image_url": { "url": "https://example.com/cup.png" },
      "frame_type": "first_frame"
    }]
  }'

Models that take a pinned frame

Not every model accepts both frame types, and the catalog is the place to check. Read supported_frame_images from GET /v1/videos/models for the model you pin. A model that does not advertise a frame_type returns a 400 instead of quietly dropping it, which is the safer behavior for a request that otherwise would cost money.

From the Video Router doc: Wan 3.0, MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 list image-to-video with an end frame, Kling 3 lists start and end frames, and Grok Imagine Video 1.5 is image-to-video only. Kling 3 does not take reference_*_urls at all, so a reference photo sent to it is an error and not a different mode.

  • Pinned opening frame: frame_images, first_frame.
  • Pinned ending: frame_images, last_frame, only if the model lists it.
  • Look or character guidance without an exact frame: input_references.

When a reference is the better choice

A reference is not a bug. If you want the product, the character or the palette from a photo but not its exact composition, such as a new camera angle on the same object, reference-to-video is the right tool. The docs say the model uses the image as visual guidance, so expect resemblance and not a pixel match.

A quick check after the clip lands: download it from the content URL and compare its first frame to your source image. If the two match, the pin worked. If the composition differs, look at the field name before you rewrite the prompt. The difference between the two modes is also covered in image-to-video vs reference-to-video, and reading the models endpoint before you pin shows how to confirm what a model accepts.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume