Grok Imagine's 4 keyframes and 7 references vs Sume's one image

xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.

4 min readSume
All posts

At xAI, Grok Imagine Video 1.5 takes up to 4 keyframes and, per a July 31, 2026 announcement, up to seven references (read 2026-10-10). Sume's grok-imagine-video-1.5 row is narrower: it needs exactly one image and rejects end frames and reference video or audio.

For anything with more than one image, route to another Sume row or call xAI directly.

What xAI offers

The xAI capabilities page (read 2026-10-10) lists five modes: text, image, reference, first and last frame, and keyframes with up to 4 images on 1.5. Reference-to-video supports up to 3 speakers with a preset voice, and the classic model's reference-to-video maximum is 10 seconds. The xAI news post (read 2026-10-10) describes image and voice references, with voice reference available on request.

These are product features of xAI's own API.

Grok Imagine multi-image modes (xAI read 2026-10-10)
ModeImagesOn Sume's Grok row
Image-to-video1Yes
First and last frame2No, end frame rejected
Keyframes (1.5)Up to 4No
Reference-to-videoUp to 7No, single reference only
Text-to-video0No, image required

What Sume's row does

The schema for grok-imagine-video-1.5 requires one image, as image_url, first_frame_url or one reference_image_urls entry, and rejects end_image_url, last_frame_url, reference video and audio, bitrate_mode, aspect_ratio and generate_audio. The registry maps only the image-to-video slot for this id.

The upside is price. The row bills a flat $0.0125 per second, so a 6-second loop is $0.075.

Sume rows for multi-image jobs

Per the Video Generation docs and the capability table, first and last frames work on Kling 3, Wan 3.0, Seedance and Omni. Multi-image references work on Wan 3.0 (10), Omni (10) and the Seedance and MiniMax rows (9). None offers xAI's preset-voice speakers; Sume's voice cloning is app-only.

A 5-second 720p Wan 3.0 clip with references bills $0.625, against $0.0625 for the same length on the Grok row.

Sume rows by image count (API code, checked 2026-10-10)
NeedSume row5 s at 720p, billed
Start and end framekling-3 (1080p only)$1.05 with audio
Start and end framewan-3.0$0.625
Up to 10 referenceswan-3.0 or Omni$0.625
Up to 9 references plus audioseedance-2.5$2.889 (text-request estimate)

Decision

If you have one still and want a cheap loop, stay on Sume's Grok row. If you want the voice features or seven references that xAI announced, use xAI. If you want multi-image work through one Sume key, use Wan 3.0 first, since it is the cheapest of the multi-image rows.

Whichever you choose, test with one short clip and check the finished job's usage.cost before you scale.

Moving an xAI multi-image workflow

Say your xAI script sends four keyframes. On Sume, the closest single-request equivalents are a start and end frame on Wan 3.0, Kling 3 or Seedance, or a set of references on Wan 3.0, which accepts up to 10 images. Neither is a keyframe sequence. If the middle frames matter, split the shot: generate frames one and two as one clip, frames two and three as the next, and join them.

That triples the request count and the cost, but each request stays inside Sume's limits. The Wan 3.0 route at 720p bills $0.625 per 5-second segment, so a three-segment sequence bills $1.875.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume