Haiku 5.5 Batch API is half price: Sume renders are not batched

Claude's Batch API gives Haiku 5.5 a 50% discount, but a Sume render is billed by Sume. How to batch the planning step and keep caps on the render.

5 min readSume
All posts

Anthropic's Batch API charges 50% less on input and output tokens, so Claude Haiku 5.5 costs $0.05 per million input tokens and $0.25 per million output tokens in batch for prompts up to 100,000 tokens. That discount applies to Anthropic tokens only. A Sume render is billed by Sume under its own meter, so a batch of Haiku planning calls can be cheap while the renders they trigger cost what they cost.

What batch covers

The pricing page says Batch and prompt caching discounts can be combined, and that Managed Agents sessions do not get the batch discount.

Haiku 5.5 batch prices per million tokens, Anthropic pricing page, read 2026-10-08
ItemStandardBatch
Input, prompt up to 100k$0.10$0.05
Output, prompt up to 100k$0.50$0.25
Input, prompt over 100k$0.50$0.25
Output, prompt over 100k$2.50$1.25

A split that uses it well

Batch is asynchronous, so use it for work that does not need an answer right away: writing 200 prompt variants, classifying briefs, drafting captions. Do the calls to Sume afterwards in a normal loop. Each Sume paid call needs its own idempotency_key, and max_spend_usd should be sent on each, because the cap applies only when present.

Do not put the paid creation inside a batch of model requests. A batch has no interactive tool loop, so there is no place to run dry_run, read the preview and then submit.

Sume's long-job pattern after planning

The wait step is described in jobs_wait for long video jobs.

  • Take a planned prompt from the batch output.
  • Call the paid tool with dry_run=true and read the preview.
  • Submit with a new idempotency_key and a max_spend_usd from your budget.
  • Wait with jobs_wait in slices of 45 to 55 seconds and never resubmit the create.
  • If the work is unattended, use Agent Completions, where generation_spend_cap_usd is required.

When batch is the wrong tool

Batch is for work with no user waiting. A job that a person is watching, such as a preview during an editing session, should use normal requests. And a loop that needs the output of one call to decide the next call, like submit then poll, cannot be batched at all because each step depends on the last.

Use batch for the wide, shallow work and normal requests for the narrow, deep loop. That keeps the discount where it applies and the caps where the spend happens.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume