Photo to video on Sume: frame_images or input_references?

Which Sume video field starts from your photo and which only guides style. If you send both, frame_images wins. Two request bodies and a check you can run.

5 min readSume
All posts

To make your photo the first frame of the video, send it in frame_images with frame_type: "first_frame"; to use it only as visual guidance, send it in input_references. If a request has both fields, frame_images controls the mode and Sume treats the call as image-to-video (Sume video docs, read 2026-10-06). One field name decides whether the photo is the opening shot or a style hint.

Teams moving off the OpenAI Sora API, which ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), often had one image field. Sume has two with different meanings, and picking the wrong one is the most common way a photo stops appearing in the output.

The two fields

Image fields on POST /v1/videos, read 2026-10-06
FieldModeWhat the model does with the imageNeeds
frame_imagesImage-to-videoUses it as the first or last frameframe_type of first_frame or last_frame
input_referencesReference-to-videoUses it as visual guidance, not an exact frameA model whose supported_input_references lists the type
Both togetherImage-to-videoframe_images winsDrop input_references to avoid confusion

Photo as the opening frame

This is the animate-my-photo case. The image URL must be public HTTPS, and the model must list first_frame in supported_frame_images.

{
  "model": "gemini-omni-flash-1.1",
  "prompt": "Slow push-in, steam rises from the cup, soft window light",
  "frame_images": [
    {
      "type": "image_url",
      "image_url": {"url": "https://example.com/cup.png"},
      "frame_type": "first_frame"
    }
  ],
  "aspect_ratio": "9:16",
  "resolution": "720p",
  "duration": 5
}

Photo as a style reference

Use input_references when you want the look, the product or the palette but not a locked first shot. The Video Router doc notes that on Omni a single reference image with no frame is reference-to-video, and that reference mode takes up to 10 images and 3 clips, each up to 3 seconds. Check supported_input_references for the model you pin: some accept image only, others add video and audio, and Kling 3 takes no reference URLs in the router table.

Because the two modes differ so much in cost and behavior, log which one each job used. A job that was meant as image-to-video but went out as reference-to-video looks fine in the queue and wrong in the output.

A check before you submit

The function reads the model's capabilities from GET /v1/videos/models and says which field to use. It needs a key and a network call, and fails loudly on an unknown model.

import os
import requests


def image_field(model: str, want_exact_first_frame: bool) -> str:
    r = requests.get(
        "https://api.sume.com/v1/videos/models",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        timeout=30,
    )
    r.raise_for_status()
    m = next(x for x in r.json()["data"] if x["id"] == model)
    if want_exact_first_frame:
        if "first_frame" not in (m.get("supported_frame_images") or []):
            raise ValueError(f"{model} has no first_frame")
        return "frame_images"
    if not m.get("supported_input_references"):
        raise ValueError(f"{model} takes no reference images")
    return "input_references"


print(image_field("gemini-omni-flash-1.1", True))

If a job ever disagrees with your intent, compare the submit body with the table above first. The existing posts on the first-frame size rule and on reference tags cover the image size and the <IMAGE_REF_0> addressing.

Which one to use for which job

A short decision rule saves a lot of confusion. If the viewer should recognize your photo as the opening shot, use frame_images: a product on a table that starts moving, a portrait that blinks. If the photo is there to tell the model what something looks like, but the shot should be new, use input_references: a bottle design that appears in a scene of your own description.

The two modes also differ in what they ask of the model. Image-to-video needs the first_frame or last_frame type in the model's supported_frame_images. Reference-to-video needs the type listed in supported_input_references. A model can support one and not the other, which is why the check function above reads both.

One last practical point: test the image URL before the batch. A photo hosted behind a login, a redirect or a short-lived link fails in a way that looks like a model problem, and the cheapest check is to open the URL in a private browser window and confirm that it loads without any cookie or header.

  • Photo as the first frame: frame_images, with frame_type: first_frame.
  • Photo as a last frame: frame_images, with frame_type: last_frame, if the model lists it.
  • Photo as a look or product guide: input_references, on a model that lists the reference type.
  • Both fields in one call: the call is image-to-video, so drop the references to keep the intent clear.
  • Always use public HTTPS URLs for images, and test one before a batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume