Batch APIs discount 50%: what a Sume bulk run changes instead

OpenAI, Anthropic and Gemini each document a 50% batch discount. Sume docs describe no batch price: a bulk queue controls throughput, not the rate.

5 min readSume
All posts

Does a Sume bulk run get a batch discount? I found none in the docs. OpenAI's Batch guide, Anthropic's batch page and the Gemini Batch API guide (all read 2026-10-04) each state a 50% discount in exchange for asynchronous turnaround. The Sume Bulk runs docs describe no batch rate: every child is an ordinary Format run, admitted through the same wallet, workspace concurrency and spend-cap checks as a single call.

That is a statement about the documentation, not about future pricing. If you need a number, read your run receipts and the billing record, not a blog post.

What a queue gives you instead

A queue gives control, not a price break. You choose concurrency from 1 to 16, the server keeps that many children in flight and refills slots as runs finish, and you set generation_spend_cap_usd per item, up to 500. The Bulk runs docs list no queue-level cap, so your worst case is the number of items times the per-item cap, and you should set a cap deliberately.

Each completed run's usage.billable_amount_usd_micros carries the generation spend attributed to it, and usage is null when it could not be read. GET /v1/usage stays the authoritative billing record.

Reading your own cost

Do not compare a discount with a run cost from memory. Run ten items, read each receipt's usage.billable_amount_usd_micros, and add them. That figure excludes the LLM turn, and the docs say debited_usd_micros is the real cost, so read both before you quote a per-video number to a finance team. usage is null, not zero, when it cannot be read; treat null as unknown.

Then compare like with like. A 50% discount on a text step that costs cents saves cents; the video step is where the spend is, and that step is outside what any of the three text batch pages describes.

Discount versus latency

Batch terms, vendor pages and Sume Bulk runs docs read 2026-10-04
OpenAI BatchAnthropic BatchesGemini BatchSume bulk run
Discount stated50%50%50%None described
Completion window24hExpires after 24 hoursTarget 24 hoursPer-run expires_at
Max submission50,000 requests100,000 requests2 GB file100 items
Your controlcustom_id mappingcustom_id mappingJob statesconcurrency 1 to 16, per-item cap

Worst-case spend before you press send

Compute the ceiling for a queue from your own caps. The numbers below are examples, not prices.

items = 100
per_item_cap_usd = 12.50     # your choice, must be 0 < cap <= 500
assert 0 < per_item_cap_usd <= 500
worst_case = items * per_item_cap_usd
print(f"worst case for {items} items: ${worst_case:,.2f}")
for c in (4, 16):
    print(f"concurrency {c}: at most {c} items in flight, "
          f"${c * per_item_cap_usd:,.2f} exposed at once")

When to use which

Use a vendor batch for text where a day is fine and the discount matters. Use a queue when each result is a media file you want back item by item with a per-run webhook, and when you want to set how many run at once. The two do not compete: copy in a batch, video in a queue is a sensible split.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume