What is Grok Imagine Video 1.5 Lite, and what does Sume run?
Grok Imagine Video 1.5 Lite is xAI's low-price video model at $0.02 a second. Sume's catalog has grok-imagine-video-1.5 instead, image-to-video only.

Grok Imagine Video 1.5 Lite is xAI's lower-cost video generation model, listed at $0.02 per second of output on xAI's models page. It takes a text prompt or a starting image and can speak lip-synced dialogue. xAI says it reaches 1080p by upscaling. Sume does not list a Lite row: the Sume video catalog has grok-imagine-video-1.5, which accepts a source image only.
What does xAI say Lite is?
xAI's video generation page names three models: grok-imagine-video-1.5, grok-imagine-video-1.5-lite and the older grok-imagine-video. It describes the 1.5 model as native 1080p and the Lite model as the cheaper option that reaches 1080p through upscaling. Clip length runs from 1 to 15 seconds with a default of 8, and audio is generated by default on both 1.5 models, with a generate_audio switch to turn it off.
xAI's own recommendation on that page is to use Lite for simple videos, because it has the lowest price per second.
- Modes on the video generation page: text-to-video, image-to-video, edit-video and extend-video. Reference images, first and last frame and keyframes are marked as 1.5-only.
- Resolutions: 480p (the default), 720p and 1080p.
- Aspect ratios listed: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9 and 5:2.
Is Lite on Sume?
No. Sume's video router documentation lists grok-imagine-video-1.5 with 480p and 720p, durations of 4 to 15 seconds, and a note that it is image-to-video only: the frame follows the aspect ratio of the source image you send. There is no text-only Grok request and no Lite id. If you send a model id that is not in the catalog, the generate route rejects it, so check the live catalog before you wire a pipeline around a model name you saw in the news.
Sume's catalog rows are listed at the provider price times 1.25, and the Grok row bills a flat 1.25 cents per output second.
How do the two compare on a 10-second clip?
The table puts the published xAI per-second prices next to the Sume catalog rate. They are different models, so this is a budget comparison and not a quality claim.
| Option | Per second | 10-second clip | Input |
|---|---|---|---|
| xAI grok-imagine-video-1.5-lite | $0.020 | $0.20 | Text or image |
| xAI grok-imagine-video-1.5 | $0.080 | $0.80 | Text or image |
| Sume grok-imagine-video-1.5 | $0.0125 | $0.125 | Image only |
Which should you pick?
If you need text-to-video from a bare prompt, Sume's Grok row cannot do it; pick a Sume model that accepts a prompt alone, or generate a still first and animate it. If you already have a product photo or key art, the image-to-video Grok row is the cheapest way on Sume to turn it into a 4 to 15 second silent clip. Submit it through the Videos endpoint, which returns a job id you poll until it completes.
Sources
Related posts
More in Models
- Whistle 16.9 MB on-device model vs Sume STT: which clips go where
Cactus Whistle transcribes 30-second clips on the device; Sume STT bills $0.01 per audio minute in the cloud. Which audio belongs on which, with a table.
- Whistle's 30-second window: passes for 5 minutes vs Sume STT
A 5-minute recording is ten 30-second Whistle passes on the device, or one Sume STT job at $0.05. How to plan chunking and when to skip it.
- Whistle keyword biasing vs Sume STT: fixing product names in captions
Whistle can bias decoding toward keywords. Sume STT takes no vocabulary list, so for captions you pass script_text instead. How the two fixes compare.
- MiMo V2.6 Pro and Flash in Sume's agent picker: what the catalog says
Sume's model catalog has rows for Xiaomi MiMo V2.6 Pro and Flash behind the OpenRouter switch. What the repo records, and what the API cannot pick.
Written by Sume