First frame plus style sheet in one video call: frame_images wins

On Sume /v1/videos, frame_images plus input_references runs as image-to-video, so plan on your style sheet being set aside. Split the call and pin the opening.

5 min readSume
All posts

If you send both frame_images and input_references in one POST /v1/videos request, Sume treats it as image-to-video: frame_images takes precedence, so the images you meant as a style sheet are not used as references. For an episode that needs an exact opening and a consistent look, make two decisions instead of one call: pin the opening with a first frame, and carry the look in the prompt or in the still itself.

This comes up as soon as a series gets an image workflow. xAI's image guide describes grok-imagine-image-2.0 generating up to 10 images per request and editing from up to 5 source images (read 2026-10-03), which makes it easy to end up with a character sheet, a location still and an approved opening all in hand. The temptation is to hand all of them to the video model at once.

Two modes, not one

The video docs describe two ways to provide images, and each triggers a different generation mode. frame_images is for image-to-video: each entry needs a frame_type of first_frame or last_frame. input_references is for reference-to-video: the model uses the images as visual guidance rather than exact frames.

The rule when both are present is plain: frame_images takes precedence and the request is treated as image-to-video. Nothing in that rule says the references are merged. Plan on them being ignored, and check the result against your expectations rather than assuming a blend.

frame_images versus input_references on /v1/videos (read 2026-10-03)
FieldModeWhat the image isCapability flag
frame_imagesImage-to-videoAn exact first or last framesupported_frame_images
input_referencesReference-to-videoStyle or content guidancesupported_input_references
Both togetherImage-to-videoReferences do not applyframe_images wins

Check what the model accepts

Both lists are published per model. GET /v1/videos/models returns supported_frame_images (for example first_frame and last_frame) and supported_input_references (image, video or audio types). A model that does not list a type will not accept it, so read the entry for the model you plan to use rather than copying a request that worked on another one.

The video docs also say audio and video references are honoured by some model families and not others, and that some accept video references but not audio. Treat that as a reason to look at the live list, not to memorise it.

import os, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
wanted = os.environ.get("VIDEO_MODEL", "seedance-2.5")

models = requests.get(f"{API}/v1/videos/models", headers=H).json()
items = models.get("data", models.get("models", []))
for m in items:
    if m.get("id") == wanted:
        print("frame images:", m.get("supported_frame_images"))
        print("input refs:  ", m.get("supported_input_references"))
        break
else:
    print("model id not in the list; check GET /v1/catalog")

What to do for an episode

Decide which job each image has. The approved opening still is a first frame. A character sheet or palette is a reference. If both matter and the model is image-to-video, bake the look into the opening still itself, which is cheap to redo, and describe the motion in the prompt. A still is a fraction of a clip's cost, so iterate there.

If the look matters more than the exact opening, drop the first frame and use references alone, then pick the best frame from the result with video frames and use it as the next episode's start. That keeps the series coherent without fighting the precedence rule.

Test once per model with a throwaway request before you queue a season. Send the same prompt with only a first frame, and again with only references, and compare. It costs two short clips and tells you whether the model keeps the face or the framing better. Use an idempotency key in each so a retry does not bill twice.

Write the choice into your episode template so nobody re-adds the second field later. A request builder that sets one of the two and refuses to set both is a few lines, and it removes a class of silent surprises.

What Sume does and does not do

Sume documents the precedence and publishes per-model support, and it runs the request you send. It does not merge a frame and a style sheet into one conditioned clip, and the docs do not say it flags references that were set aside; that is why the check belongs in your code. Vendor image limits such as the five edit sources above describe xAI's own API, not Sume's.

Sources

Related posts

More in Models

All Models posts

Written by Sume