Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0: cap Sume calls anyway
A benchmark score says how well Sonnet 5.5 finishes tasks, not what a Sume call may cost. Put generation_spend_cap_usd on every Agent Completion it starts.

No benchmark replaces a spend cap. Anthropic's Sonnet 5.5 page reports 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1 (read 2026-10-05). Those numbers describe task success. They do not limit what a tool call may spend, so the limit for a Sume call has to live in the Sume request itself.
On Sume, that limit is generation_spend_cap_usd. It is required on POST /v1/agent/completions and has no default, so a request without it fails with 400 invalid_request (Agent Completions).
What the vendor page claims, and what it leaves out
The Sonnet 5.5 page lists a model id, a price per million tokens, and agent claims such as batching tool calls more efficiently and fewer failed tool calls than Sonnet 5 (read 2026-10-05). It says nothing about the cost of the tools your agent calls. Tool spend belongs to the tool provider's meter.
| Question | Where the answer lives | Source |
|---|---|---|
| How well does the model run agent tasks? | 70.6% Terminal-Bench 4.0, 80.1% OSWorld 2.1 (vendor claims) | Anthropic page |
| What does a model token cost? | $2 input and $10 output per 1M tokens | Anthropic page |
| Does the model batch tool calls? | Vendor says more efficiently than Sonnet 5 | Anthropic page |
| What may one Sume agent run spend on generation? | generation_spend_cap_usd, required, no default | Sume docs |
| Does a paid MCP call dedupe on retry? | idempotency_key, required on write and paid tools | Sume docs |
Why a stronger agent needs the cap more
An agent that completes more tasks also makes more calls without a person watching. In the Studio chat UI, an interactive spend-approval prompt protects the wallet. A backend caller does not get that prompt, and Sume's docs say the cap replaces it. Set it per run to the most you accept for that one task.
Batched tool calls make the same point from the other side. If the model fires five paid calls in one turn, each call is its own paid submit with its own idempotency_key. The key dedupes retries. It does not limit total spend.
A start call with the cap in the body
This helper refuses to build a request without a positive cap. It then posts to the real endpoint with your key from the environment. It uses only the standard library.
import json, os, urllib.request
def start_run(instruction: str, cap_usd: float) -> dict:
if cap_usd <= 0:
raise ValueError("set a positive generation_spend_cap_usd")
body = {"instruction": instruction,
"generation_spend_cap_usd": cap_usd}
req = urllib.request.Request(
"https://api.sume.com/v1/agent/completions",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"},
method="POST")
with urllib.request.urlopen(req) as resp:
return json.load(resp)["data"]
if __name__ == "__main__":
run = start_run("Make one 6 second product teaser", 3)
print(run["id"], run["status"], run["usage"])Reading the receipt after the run
The 202 receipt already carries usage.generation_spend_cap_usd_micros, so the cap you set is echoed back before any work starts. After the run reaches a terminal status, usage records what was spent. Compare the two numbers for each task type and tighten caps that you never come close to using.
A run that returns usage: null means Sume could not read the spend, which is different from zero. Treat null as unknown and check GET /v1/usage before you assume the run was free. This habit matters more with a capable model, because capable agents finish more of the work you hand them, and a loose cap is only discovered when the invoice arrives.
What to do
Treat the vendor score as a reason to give the agent harder work, not larger budgets.
- Choose the cap per task from the metered rates on the API pricing page, as the Sume docs advise.
- Send an
Idempotency-Keyon the create call so a retried tool call returns the original receipt. - Poll the
status_urlfrom the 202 receipt, or register a webhook, instead of re-submitting. - Read
usageon the receipt after the run to compare the cap with what was spent.
Sources
Related posts
More in Agents
- Clef-flash as a yes/no gate before a paid Sume Agent Completion
Cloudflare's Clef-flash returns typed answers with probabilities. Put one in front of POST /v1/agent/completions as a filter, keep the spend cap as the guard.
- A decision model's confidence is not a spend cap: Sume's real gates
Strands Decider and Clef return confidence scores. Only the Sume gates in this table, from idempotency_key to the required spend cap, limit what a run can cost.
- DeepSeek V4.1 Flash sees images: hand one to a Sume Agent Completion
V4.1 Flash is described as natively multimodal. To act on an image with Sume, pass it as an input_image attachment on an Agent Completion, up to 30 per run.
- Does the Sume run spend cap include the LLM turn? No, only generation
On a Sume run receipt, billable_amount_usd_micros is generation spend only; the agent's LLM turn bills a separate Agent wallet. How to budget both.
Written by Sume