Grok Imagine video: cost of a 10-second clip on each tier

xAI lists grok-imagine-video-1.5 at $0.080 per second, 1.5-lite at $0.020 and the base model at $0.050. The 10-second math and a draft-then-final plan.

4 min readSume
All posts

On xAI's price list, read 2026-10-03, a 10 second clip costs $0.80 on grok-imagine-video-1.5, $0.50 on the base grok-imagine-video and $0.20 on grok-imagine-video-1.5-lite. Those totals are per-second rate times 10, nothing else.

The price list

xAI prices video per second of output.

xAI Grok Imagine pricing (read 2026-10-03)
ModelVideo price per second10 seconds
grok-imagine-video-1.5$0.080$0.80
grok-imagine-video-1.5-lite$0.020$0.20
grok-imagine-video$0.050$0.50

Images on the same page

The same xAI page also lists per-image prices of $0.02, $0.05 and $0.04 for the Grok Imagine image models. Check the page for which image model each one belongs to before you budget first frames; the video column above is per second of output.

A draft then final plan

The lite tier is a quarter of the 1.5 rate: $0.020 against $0.080 per second. If a prompt needs four tries, four lite drafts of 10 seconds cost $0.80 in total, the same as one 1.5 render. Whether a lite draft predicts the final look is something to test on your own prompts, since the vendor page makes no such claim.

On Sume, video is billed from the workspace USD balance, reserved at submit at provider list times 1.25. Where a Grok video model is in your catalog, GET /v1/videos/models returns its pricing_skus, and the poll response usage.cost is the amount billed. The Sume docs do not name Grok video ids, so confirm availability in the catalog rather than assuming it.

Checking a budget

Write the budget as clips times seconds times rate, then add a retry allowance. Fifty 10 second clips at the 1.5 rate is $40.00 before retries, and at the lite rate $10.00. Those are straight multiplications of the xAI list rates.

If one in three drafts is rejected on your own review, plan on roughly half again as many generations as finished clips. Count that from your own first batch, not from a vendor claim.

Compare before committing

Read the same three fields for any model you consider: duration range, resolution and audio support. Sume reports them as supported_durations, supported_resolutions and generate_audio. Native audio matters for cost comparisons because a model that returns synced sound can replace a separate text to speech job.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume