Gemini image batch is half price: what that means for Sume users
Google lists Gemini 3.1 Flash Image batch at $0.25 input and $30 output per 1M tokens, half standard. Sume documents no batch tier; here is how to plan.

Google's Gemini API page lists Gemini 3.1 Flash Image at $0.50 input and $60.00 output per million tokens on the standard tier and $0.25 and $30.00 on the batch tier, exactly half. Sume's docs describe no batch discount: you pay the endpoint's price line, reserved at submit, whether you submit one image or a thousand.
What does Google's batch tier change?
The page shows standard output at $60.00 per million tokens and batch output at $30.00. The per-image figures it lists (about $0.067 at 1K, $0.101 at 2K) are standard. Since batch halves the token rates, the same 1K image works out to roughly $0.034 in batch, by arithmetic on the listed tokens, not a figure printed on the page.
Batch is a trade: a lower price for work that does not need an immediate answer. Google's page is where to confirm the turnaround terms before you plan around them.
| Tier | Input per 1M tokens | Output per 1M tokens | 1K image (listed or halved) |
|---|---|---|---|
| Standard | $0.50 | $60.00 | $0.067 (listed) |
| Batch | $0.25 | $30.00 | about $0.034 (halved by arithmetic) |
Does Sume have a batch tier?
Not in the docs. Sume's pricing path for listed models is one rule: provider list price times 1.25, reserved from your wallet at submit and settled to what ran. The endpoint pricing lines are what your wallet is charged, per the image docs.
What Sume does give you for bulk work is a queue. A valid job over your concurrency limit is accepted as queued and starts when a slot opens, as Generation admission explains. That is about throughput and not about price: queued jobs bill the same as jobs that start immediately.
How should I plan a large image job?
If the work is not urgent and the price gap matters more than the plumbing, price it on Google's batch tier directly. If you want one wallet, one set of job ids, webhooks and the same API for images, video and audio, price it on Sume and pace the submits.
Pace by reading generation_limits on each submit response, and stop adding work when queue_capacity_remaining is low. A full queue returns 429 queue_full, which is not a billing event: the reservation for the failed admission is released or refunded.
curl -X POST https://api.sume.com/v1/jobs/job_123/cancel \
-H "Authorization: Bearer $SUME_API_KEY"
# cancel queued jobs you no longer need before they startWhat is the decision rule?
Take the batch price when you generate more than a few thousand images, the work can wait, and you accept running the batch yourself. Take Sume when integration time, one bill and uniform error handling are worth more than the difference. Whichever you choose, run ten images first and compare real invoices, not list figures.
Sources
Related posts
More in Pricing
- Gemini Omni Flash 1.1 at $17.50 per 1M tokens: dollars per second
Google prices Omni Flash 1.1 video output in tokens, about $0.10 a second at 720p. What that means per clip, and how Sume's per-second rate is built.
- Gemini Omni Flash 1.1 clip cost matrix: 3 to 10 seconds, 360p to 4K
A lookup table for Gemini Omni Flash 1.1: list cost per clip for 3, 5, 8 and 10 seconds at 360p, 720p, 1080p and 4K, with Sume's list x 1.25 beside it.
- Flow 4K costs 50 credits and needs Ultra; Sume lists 4K per second
Google Flow upscales Omni Flash to 1080p free for subscribers and to 4K for 50 credits on Ultra only. Sume bills 4K at $0.375 a second on any plan.
- GPT Image 2.5 edit chains: four high passes cost one max image
At 1024x1024, four high-quality GPT Image 2.5 passes add up to the output price of one max image. What that means for a one-change-per-pass edit chain on Sume.
Written by Sume