OpenRouter video: frame_images and input_references together
Send both frame_images and input_references to an OpenRouter-style video API and Sume treats it as image-to-video. What changes, and how to split the fields.

If you send frame_images and input_references in the same request to Sume's POST /v1/videos, frame_images wins: the request is treated as image-to-video and the references are not used as a separate reference-to-video mode. Sume documents this rule on its video generation page; the field names themselves come from OpenRouter's video API, which Sume follows field for field.
OpenRouter's video generation guide, read on 2026-10-03, lists both fields as optional: frame_images for image-to-video with first_frame and last_frame types, and input_references for reference-to-video style guidance. The page we fetched does not state what happens when both are present, so this post only claims what Sume does and tells you how to avoid depending on either vendor's tie-break.
What is the difference between the two fields?
They select different generation modes. frame_images pins exact pictures to the start or end of the clip. input_references gives the model visual guidance (a product, a character, a style) without fixing a frame.
Sume's docs describe each entry in frame_images as needing a frame_type of first_frame or last_frame, while input_references entries are plain image_url objects, plus video_url and audio_url on the models that accept them.
| Field | Mode it triggers | Entry shape |
|---|---|---|
| frame_images | image-to-video | image_url object plus frame_type first_frame or last_frame |
| input_references | reference-to-video | image_url object, or video_url / audio_url where the model lists them |
| both | image-to-video (frame_images take precedence) | references are not a second mode |
What happens on Sume when you send both?
The request is treated as image-to-video. Nothing in the docs says the references are blended into the frame-pinned clip, so do not rely on them to add a product or a face to a clip you have already pinned with a first frame.
This matters when you port a client that builds one request object for every model. A generic builder that always fills both arrays will silently drop one of them on Sume. If the dropped field is the one that carried your product reference, the output will look fine and be wrong.
How do you check what a model accepts?
Read GET /v1/videos/models before you build the request. Each model reports supported_frame_images (for example first_frame and last_frame) and supported_input_references (for example image_url, video_url, audio_url). Sume's docs say only models that list a type accept that type, and that audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max, while Gemini Omni Flash 1.1, higgsfield-genjutsu, and h3-max-recast accept video references but not audio.
Limits are per model, not per API. The same page says seedance-2.5 accepts 4 to 30 seconds and wan-3.0 accepts 2 to 30, so read capabilities rather than assuming one envelope.
curl "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY" \
| python3 -c "import sys,json
for m in json.load(sys.stdin)['data']:
print(m['id'], m.get('supported_frame_images'), m.get('supported_input_references'))"How should you split a request that needs both a start frame and a reference?
Pick the job each image does. If the picture must be the opening shot, use frame_images alone. If it is only a look to match, use input_references alone. If you need both on one clip, generate in two steps: create the clip from the reference first, then pull a frame and use it as the next clip's first frame.
Sume's Video Router page shows the same split under the legacy wire, with flat image_url and reference_image_urls fields. The ids are shared between the two routes, so moving between them is a path-and-body change with no id remapping.
What Sume does not do here: it has no mode that combines a pinned frame with a separate reference set in one call, and size, seed, and provider.options return 400 unsupported_parameter on this route. Check the differences table before you reuse an OpenRouter payload unchanged.
What should you do before you switch an OpenRouter client over?
Search your request builder for places that set both fields. Send each variant against GET /v1/videos/models capabilities first, then run one paid job per mode and inspect the result. Keep an Idempotency-Key on every submit so a retry returns the original job instead of creating a second one.
What is a safe test for your own client?
Before you trust a port, build three requests for one model: frame images only, references only, and both. Run all three through a dry preview where you can, then one paid job each. Inspect whether the output reflects the pinned frame, the references, or neither. Only the first two are documented modes on Sume; the third follows the precedence rule above.
Keep the requests and results next to the job ids so the next person who changes the builder can rerun them.
Sources
Related posts
More in Comparisons
- OpenRouter video webhook idempotency key vs Sume job_id
OpenRouter's video webhook sends X-OpenRouter-Idempotency-Key as job_id-status. Sume says use job_id alone. How to dedupe both, with a runnable verifier.
- OpenRouter video unsigned_urls need an API key: Sume too
OpenRouter's unsigned_urls require your API key in the Authorization header, and so does Sume's content endpoint. Why a browser video tag fails and a safe fix.
- Pocket TTS voice cloning: a wav in, and what Sume does instead
Pocket TTS clones from a wav file you pass to --voice, with consent rules in its model card. Sume's API takes voice ids, not audio. Here is the difference.
- Replicate MCP discovery via server.json vs Sume's MCP URL
Replicate publishes /.well-known/mcp/server.json for the official MCP Registry. Sume documents one hosted MCP URL and OAuth metadata. How each client connects.
Written by Sume