Luma Ray3.2 allows 16 keyframes; Sume offers first and last frame
Luma lists up to 16 keyframes per clip on its Ray3.2 API. Sume does not list Luma; its video models take a first and a last frame through frame_images.

Luma's Ray3.2 API page says Multi-Keyframe lets you set up to 16 keyframes inside a single clip. Sume does not list Luma. What Sume's video models accept is two frame positions, first_frame and last_frame, through the frame_images field of POST /v1/videos, so a 16-keyframe storyboard has to be built from several clips.
The rest of this post compares what the vendor page says with what Sume's documentation says, nothing more.
Control surface from each page
The Luma figures are from its API page read 2026-10-09; Sume's are from its video docs.
| Capability | Luma Ray3.2 (vendor page) | Sume video models |
|---|---|---|
| Keyframes in one clip | up to 16 | 2 (first_frame, last_frame) |
| Video-to-video length | up to 20 seconds | Omni edit mode via video_url; set per model |
| Output resolution | 1080p across the model | 480p to 1080p; 4K on Omni only |
| HDR and EXR | HDR generation and 16-bit EXR export | not listed |
Building a longer sequence from two-frame clips
If you need a sequence that passes through several poses, chain clips: generate clip A from frame 1 to frame 2, clip B from frame 2 to frame 3, and so on. Each clip is its own job. With Wan 3.0 on Sume at 720p, four 5-second segments cost 4 x 5 x $0.125 = $2.50, and the joins are cuts you assemble yourself.
Sume's timeline render takes video slots and renders at $0.10 per output minute with a one-minute minimum, so assembling the four segments adds $0.10 (the sheet also lists a $0.02 compose job).
When to choose which
Choose Luma if the single-clip control of many keyframes, HDR or EXR delivery is the requirement, since Sume lists none of those. Choose Sume if you want video next to audio, captions and timeline tools under one balance and two-frame control is enough. Check supported_frame_images on each model via GET /v1/videos/models before you write the request.
Sequence cost on Sume
The following figures use the same list prices as the tables above.
- A 20-second sequence from four 5-second Wan 3.0 clips at 720p is 20 x $0.125 = $2.50 for the clips.
- The same sequence on MiniMax H3 at 768p is 20 x $0.075 = $1.50, with 5 seconds being H3's minimum clip.
- Sume has no multi-keyframe field, so a plan that needs 16 poses in one clip is a reason to use Luma directly.
- Luma's page lists no per-second price in the content I retrieved, so there is no price comparison here.
Limits of this comparison
There is a second way to get more control from a two-frame model: shorten the clips. Short clips between two chosen frames give you more say over where the motion goes. At Wan 3.0's 720p rate of $0.125 per second, a 16-pose sequence at 2 seconds per pose (the model's minimum) is 15 transitions x 2 s x $0.125 = $3.75. That is a rough estimate for a plan, not a claim of equivalent output to a native 16-keyframe clip.
Sources
Related posts
More in Comparisons
- Luma Ray 3.2 reframe vs LTX reframe: 60-second 1080p, $21.60 vs $12.00
Reframing a 60-second clip to 1080p costs $21.60 on Luma's page (36 cents a second) and $12.00 on LTX's (20 cents). Lower tiers and Sume's $0.02 crop too.
- MAI-Transcribe-2 at $0.54 vs Sume STT at $0.60 an hour: the 6-cent gap
MAI-Transcribe-2-Streaming is $0.54 an audio hour through 2026; Sume STT is $0.60 an hour. At 1,000 hours the gap is $60. They suit different jobs.
- MAI-Voice-2.1 lists 23 languages; how Sume TTS sets its language
Microsoft lists 23 languages and 26 locales for MAI-Voice-2.1. Sume TTS takes one language field, defaults to English and asks you to confirm a voice mismatch.
- MAI-Voice-2.1 voice cloning is gated; how Sume TTS picks a voice
MAI-Voice-2.1 clones a voice from a 5 to 60 second clip but needs gated access. Sume TTS selects voices by avatar or voice id and has no reference-audio field.
Written by Sume