xAI Batch API discount: 20% on four text models, none on images
xAI's pricing page says the Batch API discount applies to text models only, 20% on four listed ones. Image and video batches are billed at standard rates.

Does xAI's Batch API make image or video generation cheaper? No. xAI's pricing page says the batch discount applies to text and language models only, and that image and video generation are supported in the Batch API but billed at standard rates. The discount that does exist is 20% on four named text models.
This is from xAI's pricing page, read 2026-10-02.
What does the Batch API change?
xAI's page sets the two modes side by side. Batch trades immediacy for a lower token price on some models and a lift on rate limits.
| Real-time API | Batch API | |
|---|---|---|
| Token pricing | Standard rates | Discounted rates (varies by model) |
| Response time | Immediate (seconds) | Typically within 24 hours |
| Rate limits | Per-minute limits apply | Requests do not count towards rate limits |
Which models get the discount?
The page lists a 20% discount for grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning and grok-4.20-multi-agent-0309, and says it applies to input, output, cached and reasoning tokens. Models not listed have no batch discount, and models that accept Batch without a discount show N/A on their detail page.
So for image or video work, batching buys you queueing and rate-limit headroom, not a lower price. If a workload mixes text and media, only the text calls on the four listed models earn the 20%.
Check the detail page of the exact model: the page says to toggle "Show batch API pricing" there to see its resulting batch prices.
How do I budget a big image or video run, then?
Multiply the per-unit price by the count and treat any batch as full price. For images that is count times the per-image rate; for video, seconds times the per-second rate at the resolution you choose.
Take grok-imagine-image-2.0 as an example: 1,000 images at 1K low quality are 1,000 x $0.04 = $40 whether you send them one at a time or through the Batch API. Keep a margin for retries and for edits, which add an input-image fee. The same page lists video rows per second by resolution, so a batch of 100 ten-second clips costs 1,000 seconds at the rate for the resolution you pick.
How does Sume handle a large queued run?
Sume's bulk runs queue up to 100 Format runs with a concurrency window. Each child goes through ordinary Format-run admission: wallet, workspace generation concurrency and spend caps, and an item can carry its own generation_spend_cap_usd. The bulk docs do not describe a batch price, and this post does not claim one either way; Sume's rates are on its API pricing page.
To see what a finished queue cost, sum by run_id with GET /v1/usage, as the Usage docs describe.
Sources
Related posts
More in Pricing
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
- AI video generation cost per video: what one Sume API run cost
To see what one Sume run or video job cost, call GET /v1/usage with run_id or job_id and read debited_usd: the wallet deduction, agent turns included.
- Do failed AI video generations cost credits? Reserve, capture, refund
No. Sume reserves a job's estimated USD cost at submit, captures usage only on completion, and releases or refunds the hold if the job fails or is canceled.
Written by Sume