Grok Imagine Video 1.5 reference images API: xAI vs Sume
xAI's API takes up to seven reference images, text-to-video and 1080p for Grok Imagine Video 1.5. Sume's API reference lists it as first-frame image-to-video.

Yes, the xAI API takes reference images for Grok Imagine Video 1.5, but Sume's video API does not list them for that model. On Sume, grok-imagine-video-1.5 needs a first-frame image and is not listed for reference inputs, text-only prompts or 1080p.
As of 2026-09-29, xAI's post of July 31, 2026 describes the reference feature. Sume's side comes from its Video generation docs and the Sume API reference. They describe different products: xAI's own API and Sume's catalog entry.
What did xAI add to Grok Imagine Video 1.5?
xAI's page dates the update July 31, 2026 and lists four additions. It says image references, text-to-video and native 1080p are live in the xAI API for the model id grok-imagine-video-1.5, and that voice reference support is available on request.
| Addition | What xAI's page says |
|---|---|
| Multi-reference | Each reference image locks one thing in place; up to seven references per generation |
| Text-to-video | Describe the shot, no starting image needed |
| Native 1080p | Supported with text-to-video and image-to-video |
| Voice reference | Available on request in the API |
What does Sume list for grok-imagine-video-1.5?
The API reference describes the first-frame field as "Required for grok-imagine-video-1.5", and says an end-frame image is "Not supported by grok-imagine-video-1.5". So on Sume the model starts from one image you supply.
Sume's docs name the models that honor audio and video references (the Seedance 2.x models, Wan 3.0 and MiniMax H3) and Gemini Omni Flash 1.1 for video references. Grok Imagine is not among them, and Sume's docs do not list image references for it. For reference-to-video, see reference to video by API.
Can I get 1080p from Grok Imagine on Sume?
Sume's API reference lists the ids that accept 1080p in its resolution description. grok-imagine-video-1.5 is not among them, so Sume's docs give no 1080p route for it. Check supported_resolutions from GET /v1/videos/models for the current answer.
How do I see what the catalog accepts today?
The catalog is the source of truth, and it can change. Ask it directly.
supported_resolutions: the resolutions the model accepts.supported_frame_images: which offirst_frameandlast_frameit takes.supported_input_references: which reference types it takes.
curl "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY"What should I do if I need references or 1080p?
Use a catalog model that lists what you need. Which one fits is a test on your own prompts; Kling 3 vs Grok Imagine and Seedance vs Grok Imagine compare their listed inputs. For the Grok model's own request, limits and price on Sume, see Grok Imagine video API. Sume does not say whether it will add xAI's new inputs, so build only on what the catalog lists today.
Sources
Related posts
More in Models
- Hailuo 3.0 release date: MiniMax H3 launched 2026-07-31
Hailuo 3.0 is the name people use for MiniMax H3, which MiniMax released on 2026-07-31. What the vendor pages call it, and the ids that reach it on Sume.
- How much VRAM for AI video generation? Numbers from model cards
From 14 GB to 80 GB: the VRAM figures Tencent, Wan-AI and Genmo publish for their open video models, and what each number was measured with.
- HunyuanImage 3.0 API: open weights, no listed Sume model
HunyuanImage 3.0 is an 80B-parameter open-weights image model from Tencent. Sume's image docs do not list it; here is what the card says and the hosted route.
- HunyuanVideo 1.5 low VRAM: 14 GB with offloading
Tencent's HunyuanVideo-1.5 card lists 14 GB as the minimum GPU memory with offloading on. The OOM settings and step-distilled option it names.
Written by Sume