Haiku 5.5 with US-only inference: 1.1x on both price tiers

Anthropic's pricing page applies the 1.1x US-only multiplier to Haiku 5.5's over-100k prices too. A worked agent-loop table and where Sume's own cap sits.

4 min readSume
All posts

If you pin Claude Haiku 5.5 to US-only inference with the inference_geo parameter, every price on the model gets a 1.1x multiplier, including the higher prices that apply to prompts over 100,000 tokens. Anthropic's pricing page (read 2026-10-10) states that the multiplier covers input, output, cache writes and cache reads, and says that on Haiku 5.5 it also applies to the higher over-100k prices.

That matters for an agent that drives video tools, because its prompt grows with every job result it reads. Below is the arithmetic for one request and for a 50-request loop, with the token counts labeled as assumptions.

The prices involved

From the same page (read 2026-10-10), Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Cache hits are $0.01 per million on the lower tier and $0.05 on the higher one. A request's prompt length counts all its input tokens, including cache reads and writes, and each request is priced on its own.

Global routing, the default, uses standard pricing. The page says partner platforms such as Bedrock and Google Cloud have their own regional pricing, so the 1.1x figure here is for the Claude API and the cases the page names.

One request, three prompt shapes

The token counts below are my assumptions for illustration, not measurements: a 60,000-token prompt with no cache hits, a 120,000-token prompt with no cache hits, and a 120,000-token prompt that is entirely cache reads. Each returns 1,000 output tokens.

Haiku 5.5 cost per request and per 50-request loop, assumed token counts, prices from Anthropic's pricing page (read 2026-10-10)
Prompt shapeGlobal, per requestUS-only, per requestGlobal x 50US-only x 50
60,000 input, 1,000 output$0.0065$0.00715$0.325$0.3575
120,000 input, 1,000 output$0.0625$0.06875$3.125$3.4375
120,000 cache reads, 1,000 output$0.0085$0.00935$0.425$0.4675

What the numbers say

Crossing 100,000 tokens takes the uncached request from $0.0065 to $0.0625 under my assumptions, nearly ten times as much for twice the prompt. The cache-hit row shows why a stable prefix matters: the same 120,000 tokens as cache reads cost $0.0085, because reads are billed at $0.05 per million even on the higher tier.

US-only adds a flat 10 percent on top of whichever row you are in. It does not change which tier you are in.

  • Keep tool results short. One batch jobs_wait with include_results brings each result into the prompt once, instead of a wait plus a separate read per job.
  • Do not paste media URLs or full result objects back into the prompt if the next step needs only the job id.
  • Compact or restart the agent session before it nears 100,000 tokens if the work allows.

Where Sume's own cost sits

None of the numbers above are Sume charges. If your agent runs in your own client, you pay Anthropic for tokens and Sume for the generations it submits. Over hosted MCP, the Sume side is guarded by idempotency_key, optional dry_run and max_spend_usd on paid tools, and changing the model behind the client does not change those gates.

If you call Sume's own agent instead, Agent Completions accept only sume-agent as the model and require generation_spend_cap_usd on every request. Sume does not let you choose Haiku 5.5 there, so the inference_geo multiplier is a question for the clients you run yourself.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume