Image-to-video not starting on my photo: frame_images vs references
Your photo is a reference, not a first frame, when it goes in input_references. Use frame_images with first_frame on Sume /v1/videos to pin the opening shot.

If your image-to-video clip does not open on your photo, the photo almost certainly went into the wrong field. On Sume POST /v1/videos, an image in frame_images with frame_type: "first_frame" is pinned as the opening frame. The same image in input_references is treated as visual guidance for reference-to-video, so the model is free to compose a different opening shot that only resembles it.
Sume infers the mode from which field is present and never asks you to declare it. That is convenient, but it also means a field mix-up does not return an error: you get a valid clip that just does not start where you expected. The rest of this page is a short way to tell which mode you actually requested, and how to fix the request.
Which field starts which mode
The video generation guide describes two ways to send images. The mode follows the field, so read the request you sent against this table before you blame the model.
| You sent | Mode the API infers | What the image does |
|---|---|---|
| frame_images with first_frame | Image-to-video | Pinned as the opening frame |
| frame_images with last_frame | Image-to-video | Pinned as the closing frame, on models that list last_frame |
| input_references (image_url) | Reference-to-video | Style or content guidance, not an exact frame |
| Both fields in one body | Image-to-video | frame_images wins, input_references are not the mode |
| Neither | Text-to-video | No image used |
Fix the request
Move the photo out of input_references and into frame_images. Here is a minimal request that pins one photo as the first frame of a Seedance 2.0 clip. Replace the URL with a publicly reachable HTTPS image, because the API downloads it. If the image is behind a login, a signed link that has expired, or a private bucket, the job fails with a message about an input media URL that could not be downloaded, so test the link in a private browser window first.
Two smaller points save a retry. First, aspect ratio: the pinned frame sets the composition, so choose an aspect_ratio that the model lists and that matches your image, instead of expecting the model to crop it for you. Second, one frame is enough for a first-frame pin. Add a last_frame entry only when the model lists it and you want to control the ending as well.
The legacy Video Router surface follows the same rule with flat fields. Its Gemini Omni Flash 1.1 table says that a single reference image with no frame is reference-to-video, which is the same rule that was applied to Seedance after issue 2044. If you use /v1/video-router/generate, the field to use for a pinned opening frame is image_url, not reference_image_urls.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "The camera slowly pushes in as steam rises from the cup",
"resolution": "720p",
"frame_images": [{
"type": "image_url",
"image_url": { "url": "https://example.com/cup.png" },
"frame_type": "first_frame"
}]
}'Models that take a pinned frame
Not every model accepts both frame types, and the catalog is the place to check. Read supported_frame_images from GET /v1/videos/models for the model you pin. A model that does not advertise a frame_type returns a 400 instead of quietly dropping it, which is the safer behavior for a request that otherwise would cost money.
From the Video Router doc: Wan 3.0, MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 list image-to-video with an end frame, Kling 3 lists start and end frames, and Grok Imagine Video 1.5 is image-to-video only. Kling 3 does not take reference_*_urls at all, so a reference photo sent to it is an error and not a different mode.
- Pinned opening frame:
frame_images,first_frame. - Pinned ending:
frame_images,last_frame, only if the model lists it. - Look or character guidance without an exact frame:
input_references.
When a reference is the better choice
A reference is not a bug. If you want the product, the character or the palette from a photo but not its exact composition, such as a new camera angle on the same object, reference-to-video is the right tool. The docs say the model uses the image as visual guidance, so expect resemblance and not a pixel match.
A quick check after the clip lands: download it from the content URL and compare its first frame to your source image. If the two match, the pin worked. If the composition differs, look at the field name before you rewrite the prompt. The difference between the two modes is also covered in image-to-video vs reference-to-video, and reading the models endpoint before you pin shows how to confirm what a model accepts.
Sources
Related posts
More in Developers
- Japanese speech to text API: Sume STT with language_code ja
Transcribe Japanese audio with Sume STT: send language_code ja, read word times, and test a sample first. $0.01 per audio minute, 10 minute jobs.
- Job id or run id? Which Sume endpoint to poll for each product
Jobs, Format runs, Actions and Agent Completions have different ids, poll URLs and webhook events. Which to poll for each product, and which SDK helper to call.
- Let browsers start Sume jobs through your server, not with your key
Browsers must never hold a Sume API key. A server route authenticates the user, checks the input, derives an Idempotency-Key and returns only the status URL.
- How do I add a listen-to-this-page audio version with TTS?
Turn each article into an audio file with one async TTS job per page: a 9,000-character article costs 43 cents on Sume. What it does not replace.
Written by Sume