One image to Seedance: reference, or first frame?

On Sume a single reference image with no frame field is priced and routed as reference-to-video; add a first frame to get image-to-video. How to choose.

5 min readSume
All posts

If you send one image to Seedance on Sume without saying it is a frame, the request is treated as reference-to-video: the image guides the content but is not the first frame of the clip. To make that image the opening frame, send it as frame_images with frame_type: "first_frame" on /v1/videos, or as image_url on the Video Router surface.

The two modes cost the same per second of output, but they behave differently, so picking the wrong one is the usual reason a Seedance clip "does not start from my picture".

How does Sume decide which mode a request is?

Sume picks the provider endpoint from the shape of the request, and prices the same endpoint. The rule in its pricing code mirrors the payload builder: a reference video, a reference audio clip, more than one reference image, or exactly one reference image with no frame field all go to reference-to-video. A frame field alone goes to image-to-video; nothing goes to text-to-video.

This table describes how Sume's pricing code picks the endpoint it estimates against, and the code comments say it mirrors the Seedance payload builder. The sure way to check what your own request did is the job's model and cost on the poll response, then a one-image test in each shape at 480p.

Seedance request shape and the mode it gets on Sume, read 2026-10-02
What you sendMode
prompt onlyText-to-video
first frame (and optional last frame)Image-to-video
one reference image, no frameReference-to-video
two or more reference imagesReference-to-video
any reference video or audioReference-to-video
frame_images plus input_references on /v1/videosImage-to-video; input_references are dropped

What is the difference in the output?

Image-to-video starts from your image: frame 1 is that picture, then it moves. Reference-to-video treats the image as an identity or style input. The model may place your product in a new composition, a new angle or a new setting, which is what you want for a character who appears in several shots, and not what you want for an animated still.

Prompts differ too. In reference mode you name the image in the prompt with a tag, as covered in the @image tag post; in frame mode you describe motion from the picture, as in the first and last frame post.

Which should you use for a product shot?

Use the frame when the packshot must look exactly like the photo at the start, such as a label that has to stay readable. Use a reference when the product should appear in a scene you describe and exact framing is not required. Do not expect one request to do both: a frame is the exact opening picture, and references are guidance, which is why the Video generation docs make the frame win when both are sent.

On /v1/videos, sending both frame_images and input_references makes the request image-to-video and drops the references; the which-wins post goes through it. The mode is decided by the shape of the request, and a frame always wins on that surface.

How do you send each one?

The two request bodies on /v1/videos differ in one field. Both use a public HTTPS URL.

// Image-to-video: the picture is frame 1
{
  "model": "seedance-2.5",
  "prompt": "Slow push-in, steam rises from the cup",
  "frame_images": [{
    "type": "image_url",
    "image_url": {"url": "https://example.com/cup.png"},
    "frame_type": "first_frame"
  }],
  "duration": 8, "resolution": "720p", "aspect_ratio": "16:9"
}

// Reference-to-video: the picture guides the content
{
  "model": "seedance-2.5",
  "prompt": "The cup from @Image 1 on a cafe table, morning light",
  "input_references": [{
    "type": "image_url",
    "image_url": {"url": "https://example.com/cup.png"}
  }],
  "duration": 8, "resolution": "720p", "aspect_ratio": "16:9"
}

Does the mode change the price?

Not through the image. Sume's estimate depends on resolution, aspect ratio, output seconds and the presence of a reference video; an image in either role adds nothing, as the price-effect post shows. The two bodies above estimate at the same amount. What changes is the provider endpoint that runs the job, and with it how closely the first frame matches your picture.

If a clip needs the picture as frame 1 and also needs identity to hold for the next 15 seconds, generate the opening shot from the frame and use a reference-to-video clip for the rest, then join them. See AI video from multiple images for the options.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume