Ray 3.2 takes 16 keyframes; Sume frame_images takes two
Luma lists up to 16 keyframes per Ray 3.2 clip. Sume's frame_images array uses first_frame and last_frame. Which catalog models list them, per the docs.

Luma lists up to 16 keyframes per clip for Ray 3.2. Sume does not run Ray 3.2, and its frame_images field accepts two frame types, first_frame and last_frame, on models that list them in supported_frame_images.
That gives you a start and an end anchor per clip. For anything more, chain clips.
How frame_images works on Sume
In the Video generation docs, frame_images carries first and last frames for image-to-video. Each entry has an image_url object and a frame_type. If you send both frame_images and input_references, frame_images controls the mode and Sume runs image-to-video.
The Video Router takes the flat fields instead: image_url and an optional end_image_url for Gemini Omni Flash 1.1, and the docs list start/end-frame image-to-video for minimax-h3-max.
Which ids list which anchors
| Catalog id | First frame | Last frame | Source in docs |
|---|---|---|---|
| seedance-2 | yes | yes | Video Models example, supported_frame_images |
| gemini-omni-flash-1.1 | image_url | end_image_url | Video Router capability table |
| minimax-h3-max | yes | yes | start/end-frame image-to-video |
| Others | read supported_frame_images | read supported_frame_images | GET /v1/videos/models |
Request shape
A first and last frame request on /v1/videos carries two entries in frame_images. Send the model's supported durations and resolutions, as listed by GET /v1/videos/models.
{
"model": "seedance-2",
"prompt": "The door opens and she walks into the light",
"duration": 6,
"frame_images": [
{"type": "image_url", "image_url": {"url": "https://example.com/open.png"}, "frame_type": "first_frame"},
{"type": "image_url", "image_url": {"url": "https://example.com/close.png"}, "frame_type": "last_frame"}
]
}Getting from two anchors to sixteen
Treat each pair of adjacent keyframes as one clip: keyframe 1 as first_frame and keyframe 2 as last_frame, then 2 to 3, and so on. Fifteen clips cover sixteen keyframes. Each clip is its own job with its own reserve, and the joins are the visible risk: a clip that ends on keyframe 2 and the next one that starts on it will share one identical frame, so trim a frame at each join or hide it with a short fade in Timeline 1.0.
This is more expensive and slower than one 16-keyframe generation, and it cannot give the model a global view of all sixteen anchors. The trade is control per segment.
Limits to respect
Duration is per model: 4 to 15 seconds for the seedance-2 example, 5 to 15 for minimax-h3-max, and 3 to 10 for Omni Flash 1.1. Remember that Sume pins one model for each request, so keep the model constant across a chain to avoid a visible change in look. Sume bills the provider list times 1.25 for each job.
A 4-keyframe worked example
Say you have keyframes K1 to K4. Run three jobs: K1 to K2, K2 to K3 and K3 to K4, each with first_frame and last_frame. Three jobs of 5 seconds give a 15 second sequence on seedance-2, whose catalog entry lists first and last frames. Join the three clips in Timeline with 0.25 second fades or hard cuts.
With 16 keyframes the same pattern needs 15 jobs, 75 seconds of footage at 5 seconds each, and a render of ceil(75/60) = 2 minutes at $0.10 per minute, so $0.20 for the join.
What you lose
A single 16-keyframe generation can plan motion across all anchors. A chain of pairs cannot, and each pair is sampled independently. The visible symptom is a small change of lighting or detail at each keyframe. Fix it by choosing keyframes with matching lighting and by using a short crossfade at the join.
Verify live
The table above reflects the docs on the date shown. Catalogs change, so before you build, call GET /v1/videos/models and read supported_frame_images for the id you want. An id that lists only first_frame cannot take an end image, and sending last_frame to it would not work. Reference images are a separate mode and do not act as exact frames.
Before you build
Before you build, read the linked Sume docs page for the exact request fields, limits and prices, because those pages are the source of truth and can change. Run one short, cheap test with your own material first, check the output in a player and in your editor, and only then scale to the full shot list. Keep every job id and file you approve, so a later change never forces you to regenerate work that was already signed off. Note that this post describes Sume's catalog and tools; Sume does not run Luma Ray 3.2, and nothing here claims HDR or EXR output.
Sources
Related posts
More in Comparisons
- Ray 3.2 tracks 8 faces; Sume swaps 1-4 people with H3 Max Recast
Luma Ray 3.2 lists facial tracking for up to 8 faces. Sume has no tracking output, but h3-max-recast swaps 1-4 people in a clip and Kling drives a still.
- Reference limits per request, vendor vs Sume: Wan 3.0, Seedance, Omni
What each vendor says one video request can take, next to what Sume's catalog allows: Wan 3.0, Seedance 2.5, Gemini Omni 1.1 Flash, and Ray3.2.
- Replicate's default webhook secret versus Sume's signing secret
Replicate's secret is at /v1/webhooks/default/secret and it signs id.timestamp.body. Sume's is at /v1/webhooks/signing-secret and signs timestamp.raw_body.
- Retouch 200 product photos: Image API, Agent Completion or Format?
For 200 identical retouches use POST /v1/images: fixed model, price per image, plan queue limits. Use an Agent Completion only if each photo needs judgement.
Written by Sume