Photo to video on Sume: frame_images or input_references?
Which Sume video field starts from your photo and which only guides style. If you send both, frame_images wins. Two request bodies and a check you can run.

To make your photo the first frame of the video, send it in frame_images with frame_type: "first_frame"; to use it only as visual guidance, send it in input_references. If a request has both fields, frame_images controls the mode and Sume treats the call as image-to-video (Sume video docs, read 2026-10-06). One field name decides whether the photo is the opening shot or a style hint.
Teams moving off the OpenAI Sora API, which ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), often had one image field. Sume has two with different meanings, and picking the wrong one is the most common way a photo stops appearing in the output.
The two fields
| Field | Mode | What the model does with the image | Needs |
|---|---|---|---|
| frame_images | Image-to-video | Uses it as the first or last frame | frame_type of first_frame or last_frame |
| input_references | Reference-to-video | Uses it as visual guidance, not an exact frame | A model whose supported_input_references lists the type |
| Both together | Image-to-video | frame_images wins | Drop input_references to avoid confusion |
Photo as the opening frame
This is the animate-my-photo case. The image URL must be public HTTPS, and the model must list first_frame in supported_frame_images.
{
"model": "gemini-omni-flash-1.1",
"prompt": "Slow push-in, steam rises from the cup, soft window light",
"frame_images": [
{
"type": "image_url",
"image_url": {"url": "https://example.com/cup.png"},
"frame_type": "first_frame"
}
],
"aspect_ratio": "9:16",
"resolution": "720p",
"duration": 5
}Photo as a style reference
Use input_references when you want the look, the product or the palette but not a locked first shot. The Video Router doc notes that on Omni a single reference image with no frame is reference-to-video, and that reference mode takes up to 10 images and 3 clips, each up to 3 seconds. Check supported_input_references for the model you pin: some accept image only, others add video and audio, and Kling 3 takes no reference URLs in the router table.
Because the two modes differ so much in cost and behavior, log which one each job used. A job that was meant as image-to-video but went out as reference-to-video looks fine in the queue and wrong in the output.
A check before you submit
The function reads the model's capabilities from GET /v1/videos/models and says which field to use. It needs a key and a network call, and fails loudly on an unknown model.
import os
import requests
def image_field(model: str, want_exact_first_frame: bool) -> str:
r = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
m = next(x for x in r.json()["data"] if x["id"] == model)
if want_exact_first_frame:
if "first_frame" not in (m.get("supported_frame_images") or []):
raise ValueError(f"{model} has no first_frame")
return "frame_images"
if not m.get("supported_input_references"):
raise ValueError(f"{model} takes no reference images")
return "input_references"
print(image_field("gemini-omni-flash-1.1", True))If a job ever disagrees with your intent, compare the submit body with the table above first. The existing posts on the first-frame size rule and on reference tags cover the image size and the <IMAGE_REF_0> addressing.
Which one to use for which job
A short decision rule saves a lot of confusion. If the viewer should recognize your photo as the opening shot, use frame_images: a product on a table that starts moving, a portrait that blinks. If the photo is there to tell the model what something looks like, but the shot should be new, use input_references: a bottle design that appears in a scene of your own description.
The two modes also differ in what they ask of the model. Image-to-video needs the first_frame or last_frame type in the model's supported_frame_images. Reference-to-video needs the type listed in supported_input_references. A model can support one and not the other, which is why the check function above reads both.
One last practical point: test the image URL before the batch. A photo hosted behind a login, a redirect or a short-lived link fails in a way that looks like a model problem, and the cheapest check is to open the URL in a private browser window and confirm that it loads without any cookie or header.
- Photo as the first frame:
frame_images, withframe_type: first_frame. - Photo as a last frame:
frame_images, withframe_type: last_frame, if the model lists it. - Photo as a look or product guide:
input_references, on a model that lists the reference type. - Both fields in one call: the call is image-to-video, so drop the references to keep the intent clear.
- Always use public HTTPS URLs for images, and test one before a batch.
Sources
Related posts
More in Developers
- Score a Sora replacement: five numbers to log for every clip
Before you cut over from Sora to a Sume video model, log five numbers per clip: status, wall time, usage.cost, duration asked and failures. Script included.
- Prove a Sume video retry is safe: same key, same job, one charge
A runnable test for a ported Sora worker: submit twice with one Idempotency-Key on /v1/videos, assert the same job id came back, and cancel before it bills.
- Sume job statuses: three vocabularies and a Python normalizer
Ported Sora code checks one status word. Sume has pending, queued and IN_QUEUE depending on the route. A table, and a normalizer you can run with no network.
- Timeout for AI video jobs: set a deadline in your worker, not HTTP
A Sora-era HTTP timeout of 10 minutes breaks on Sume: sync waits cap at 30 s while jobs run for minutes. Use a client deadline and a poll. Python example.
Written by Sume