frame_images beats input_references on Sume /v1/videos

If one /v1/videos request has both frame_images and input_references, Sume runs image-to-video and uses the frames. How to keep a style reference working.

4 min readSume
All posts

On Sume's /v1/videos, if a request carries both frame_images and input_references, frame_images controls the mode and Sume processes the job as image-to-video. Your style references are not used as references in that request. If you wanted both a fixed first frame and a style reference, check what your model supports before you send both.

The two fields

frame_images vs input_references (Sume docs, read 2026-10-05)
FieldModeWhat the model does with the image
frame_images (first_frame, last_frame)image-to-videoUses the image as an exact frame
input_referencesreference-to-videoUses the image as visual guidance, not an exact frame

The trap

The docs describe the precedence in one sentence, so a request with both fields does not return an error. It runs in the other mode. The symptom is a clip that starts with your first frame and ignores the look of your reference image. You pay for a clip that did what the docs say it does, not what you meant.

{
  "model": "seedance-2",
  "prompt": "The character walks through a forest",
  "frame_images": [{
    "type": "image_url",
    "image_url": {"url": "https://example.com/first-frame.png"},
    "frame_type": "first_frame"
  }],
  "input_references": [{
    "type": "image_url",
    "image_url": {"url": "https://example.com/style-ref.png"}
  }]
}

What to do instead

Decide the mode per clip. If you need an exact opening frame, send only frame_images. If you need the look of a reference, send only input_references. On the Video Router wire, the same split is image_url for a frame and reference_image_urls for references. The docs say a model accepts a reference type only if its supported_input_references includes it, so read that list: Seedance 2.x, Wan 3.0 and the MiniMax H3 models accept audio and video references, while Gemini Omni Flash 1.1 accepts video references but not audio.

Why the docs make this choice

A first frame and a set of references are two different instructions. A first frame says what the clip starts as. References say what the clip should look like. When both are present Sume has to pick one mode, and the docs say frame_images controls it. A client that sends both because its form had two upload boxes will see the frame honored and the references ignored, with no error.

In your own client, guard against it: if frame_images is non-empty, drop input_references before you send, or show the user a message that the two cannot be combined. That makes the result match what the user sees on screen.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

If you are new to the API, start with one clip, one model and the lowest resolution, read the full response once, and only then build a loop around it. Most surprises with video jobs come from fields that were defaulted, not from fields that were set.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume