Haiku 5.5 Batch API is half price: Sume renders are not batched
Claude's Batch API gives Haiku 5.5 a 50% discount, but a Sume render is billed by Sume. How to batch the planning step and keep caps on the render.

Anthropic's Batch API charges 50% less on input and output tokens, so Claude Haiku 5.5 costs $0.05 per million input tokens and $0.25 per million output tokens in batch for prompts up to 100,000 tokens. That discount applies to Anthropic tokens only. A Sume render is billed by Sume under its own meter, so a batch of Haiku planning calls can be cheap while the renders they trigger cost what they cost.
What batch covers
The pricing page says Batch and prompt caching discounts can be combined, and that Managed Agents sessions do not get the batch discount.
| Item | Standard | Batch |
|---|---|---|
| Input, prompt up to 100k | $0.10 | $0.05 |
| Output, prompt up to 100k | $0.50 | $0.25 |
| Input, prompt over 100k | $0.50 | $0.25 |
| Output, prompt over 100k | $2.50 | $1.25 |
A split that uses it well
Batch is asynchronous, so use it for work that does not need an answer right away: writing 200 prompt variants, classifying briefs, drafting captions. Do the calls to Sume afterwards in a normal loop. Each Sume paid call needs its own idempotency_key, and max_spend_usd should be sent on each, because the cap applies only when present.
Do not put the paid creation inside a batch of model requests. A batch has no interactive tool loop, so there is no place to run dry_run, read the preview and then submit.
Sume's long-job pattern after planning
The wait step is described in jobs_wait for long video jobs.
- Take a planned prompt from the batch output.
- Call the paid tool with
dry_run=trueand read the preview. - Submit with a new
idempotency_keyand amax_spend_usdfrom your budget. - Wait with
jobs_waitin slices of 45 to 55 seconds and never resubmit the create. - If the work is unattended, use Agent Completions, where
generation_spend_cap_usdis required.
When batch is the wrong tool
Batch is for work with no user waiting. A job that a person is watching, such as a preview during an editing session, should use normal requests. And a loop that needs the output of one call to decide the next call, like submit then poll, cannot be batched at all because each step depends on the last.
Use batch for the wide, shallow work and normal requests for the narrow, deep loop. That keeps the discount where it applies and the caps where the spend happens.
Sources
Related posts
More in Pricing
- Higgsfield Genjutsu at 30 seconds: $11.93 at 480p, $25.54 at 720p
The longest source video Genjutsu accepts on Sume is 30 seconds. Price at 480p and 720p, a per-second rate, a Wan 3.0 comparison and the request shape.
- How AI avatar tools bill: credits, minutes, or seconds (Oct 2026)
Synthesia and HeyGen bill credits, Tavus bills conversation minutes with a 30-second floor, LiveAvatar bills credits per 30 seconds. Sume bills seconds.
- How many 30-second Wan 3.0 clips does $25, $100 or $250 buy?
Whole 30-second clips per balance at 480p, 720p and 1080p on Sume's wan-3.0 pricing, plus what is left over, using the flat per-second rate.
- How many Gemini Omni Flash clips does $25 buy on Sume?
A $25 balance buys 40 five-second Omni Flash clips at 720p, 26 at 1080p, 13 at 4K or 133 at 360p on Sume. Counts for 3, 5 and 10 seconds at every resolution.
Written by Sume