xAI recommends Grok Imagine Video 1.5 ($0.08/s) over the $0.05 model
xAI lists grok-imagine-video-1.5 at $0.080 per second and grok-imagine-video at $0.050, and recommends 1.5. Sume carries only 1.5, image-to-video.

xAI lists grok-imagine-video-1.5 at $0.080 per second and grok-imagine-video at $0.050 per second, and its docs recommend Grok Imagine Video 1.5. So the 5 cent model is the cheaper one, but it is not the one xAI points new work to. Sume's catalog carries only grok-imagine-video-1.5, as an image-to-video model.
The two xAI models
Prices are from the xAI models page. The totals are my arithmetic, rate x seconds.
| Clip | grok-imagine-video-1.5 ($0.080/s) | grok-imagine-video ($0.050/s) |
|---|---|---|
| 4 s | $0.32 | $0.20 |
| 5 s | $0.40 | $0.25 |
| 10 s | $0.80 | $0.50 |
| 15 s | $1.20 | $0.75 |
When the cheaper model still fits
1.5 costs 60% more per second ($0.080 / $0.050 = 1.6). The page I read does not say how the two models differ in output, so the reason to pay the premium is whatever you see in your own test. The xAI guide says that duration and resolution both affect the total cost, and that video can be up to 15 seconds.
How to decide
Pick by what you can verify. Start from xAI's recommendation, 1.5, and test the cheaper model only if the cost gap matters at your volume. At 1,000 clips of 10 seconds, the gap is 1,000 x 10 x ($0.080 - $0.050) = $300 at xAI's listed rates. If the cheaper model's output meets your bar, that is real money. If it does not, 1.5 is the model xAI points you to.
- Run the same 5 prompts through both models and compare side by side.
- Check the length: the xAI guide says video can run up to 15 seconds, and it names only 1.5 as the model id.
- On Sume, the choice is already made: only 1.5 is listed.
What Sume lists
Sume's Video Router lists grok-imagine-video-1.5 at 480p and 720p, for 4 to 15 seconds, as image-to-video only. It has no text-to-video mode and no reference input on Sume. Sume lists no grok-imagine-video. The Sume price per second is on GET /v1/videos/models, and I do not state it here.
xAI's guide says video generation starts from one source still image. That matches the Sume row.
| Item | xAI | Sume |
|---|---|---|
| Model ids | grok-imagine-video-1.5, grok-imagine-video | grok-imagine-video-1.5 only |
| Mode | Starts from one source still image | Image-to-video only |
| Length | Up to 15 s | 4 to 15 s |
| Resolutions | Not in the pages I read | 480p, 720p |
Sources
Related posts
More in Models
- grok-imagine-video-1.5 on Sume: needs an image, seven fields refused
grok-imagine-video-1.5 is image-to-video only on Sume: send one image, no end frame, no reference video or audio, no aspect_ratio, no generate_audio.
- H3 Max 1080p regenerates from 768p: the price of the extra step
MiniMax says its 2K path regenerates in context. Sume documents H3 Max 1080p as a latent refinement of 768p, $0.20 vs $0.10 per second. When the doubling pays.
- H3 or H3 Max? Pick the Sume row by job, with per-clip prices
Both read references and make stereo audio. H3 is $0.075 per second at 768p, H3 Max $0.10 with a 1080p row. A job-by-job guide with 10-second prices.
- higgsfield-genjutsu missing from /v1/videos/models: why
higgsfield-genjutsu is in the Sume video catalog only when its provider is configured. It is Motion Transfer: one video_url plus 1-8 images, 480p or 720p.
Written by Sume