Grok Imagine video in Sume: why it needs a start-frame image
Sume's Grok Imagine entry is image-to-video only: it blocks a submit without a start frame, tops out at 10 seconds, and sends no audio or aspect ratio.

Grok Imagine in Sume's Agents Videos panel needs a start-frame image: submit without one and the panel stops with the message "Grok Imagine requires a start-frame image." It is image-to-video only, offers 5, 6, 8 or 10 seconds at 480p or 720p, and sends neither audio nor aspect ratio.
The limit is in Sume's panel, not in xAI's model. xAI's video generation guide lists grok-imagine-video-1.5, describes both text-to-video and animating a still image, and says duration can run up to 15 seconds. Sume's panel wires only the image-to-video path and stops at 10 seconds.
What does the panel send for Grok Imagine?
When you pick Grok Imagine, the panel maps it to the grok-imagine-video-1.5 model id and builds a smaller request than for other models. Aspect ratio is left out, audio is left out, and no end image is sent. Only the prompt, duration, resolution and the start-frame URL travel.
That matters if you are used to setting a 9:16 aspect ratio on other models. For Grok, prepare the start frame in the shape you want, because the panel will not ask the model for a ratio.
| Control | Sume panel | xAI guide |
|---|---|---|
| Start image | Required | Optional (text-to-video also described) |
| Duration | 5, 6, 8 or 10 s | Up to 15 s |
| Resolution | 480p or 720p | Not specified on the page |
| Audio toggle | None | Not specified on the page |
| End frame | No | Not specified on the page |
Can I do text-to-video with Grok Imagine on Sume?
Not through this panel, even though xAI describes text-to-video. Without a start frame the submit is blocked before it reaches the API. If you need text-only prompts, pick Auto, Wan 3.0, MiniMax H3 or Kling 3.0, which accept an optional start frame rather than requiring one.
The existing post on Grok Imagine 1.5 text-to-video versus Sume's image-to-video-only row goes further into the vendor side.
When is Grok Imagine the right pick?
Pick it when you already have a strong still, such as a product shot or a character frame, and want expressive motion from it for a short clip. For a hero image that must be preserved, an image-to-video model with a start frame keeps the first frame anchored.
Pick something else when you need audio, a fixed last frame, references, or more than 10 seconds. Sume's docs say every catalog model other than the long-clip rows tops out at 15 seconds, and several of the longer ones are covered in which AI video model for a 15-second single take.
- Start frame in hand, short clip, no audio needed: Grok Imagine.
- Need audio: Wan 3.0, MiniMax H3 or H3 Max.
- Need a last frame: Wan 3.0 or MiniMax H3.
- Need to run from text only: not Grok Imagine.
How does it compare with the other panel models?
Within the same panel, Kling 3.0 offers 5 to 15 seconds up to 4K, Wan 3.0 runs 2 to 30 seconds, and MiniMax H3 runs 5 to 15 seconds at native 768p with stereo audio. Grok Imagine has the lowest top resolution and a 10-second ceiling in that set, which is why it suits quick motion tests from a still rather than final delivery.
A practical pattern is to test a still with Grok Imagine at 480p, decide whether the motion idea works, then move to a model with a last frame or audio for the keeper. See Seedance versus Grok Imagine for the other direction of the comparison.
Sume does not publish a benchmark ranking between these models and this post does not either; the table above is a control comparison, not a quality one.
What should I check before submitting?
Check that the start-frame URL is reachable. The API takes public HTTPS media URLs, and the docs list inaccessible reference images as a first thing to check when a generation fails. See the video generation docs for the troubleshooting list.
Billing follows the rest of the video catalog: Sume reserves provider list times 1.25 on submit, and the poll response's usage.cost is the billable amount. Read it from the job instead of estimating from the panel.
Sources
Related posts
More in Models
- H3 Max 3D to Video: previs to photoreal on fal, not on Sume
fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.
- H3 Max Insert-Video: add a scene mid-clip on fal, and on Sume
fal's H3 Max Insert-Video adds 5 to 13 s to a source up to 60 s, billed on the new seconds only. Sume lists no such row; here is a trim, generate, join route.
- H3 Max Recast prompt: optional, and what Sume does with one
On Sume, H3 Max Recast runs without a prompt: the 1-4 photos say who to swap in. A prompt is allowed up to 2000 characters. Request body and limits below.
- Kling 3.0 4K in Sume's Videos panel: what actually gets submitted
The panel lists 4K for Kling 3.0 but its live submit maps 4K down to 1080p. Where 4K does exist on Sume's API, and how to check the model before you pay.
Written by Sume