Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0: cap Sume calls anyway

A benchmark score says how well Sonnet 5.5 finishes tasks, not what a Sume call may cost. Put generation_spend_cap_usd on every Agent Completion it starts.

5 min readSume
All posts

No benchmark replaces a spend cap. Anthropic's Sonnet 5.5 page reports 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1 (read 2026-10-05). Those numbers describe task success. They do not limit what a tool call may spend, so the limit for a Sume call has to live in the Sume request itself.

On Sume, that limit is generation_spend_cap_usd. It is required on POST /v1/agent/completions and has no default, so a request without it fails with 400 invalid_request (Agent Completions).

What the vendor page claims, and what it leaves out

The Sonnet 5.5 page lists a model id, a price per million tokens, and agent claims such as batching tool calls more efficiently and fewer failed tool calls than Sonnet 5 (read 2026-10-05). It says nothing about the cost of the tools your agent calls. Tool spend belongs to the tool provider's meter.

Sonnet 5.5 vendor claims vs Sume controls, read 2026-10-05
QuestionWhere the answer livesSource
How well does the model run agent tasks?70.6% Terminal-Bench 4.0, 80.1% OSWorld 2.1 (vendor claims)Anthropic page
What does a model token cost?$2 input and $10 output per 1M tokensAnthropic page
Does the model batch tool calls?Vendor says more efficiently than Sonnet 5Anthropic page
What may one Sume agent run spend on generation?generation_spend_cap_usd, required, no defaultSume docs
Does a paid MCP call dedupe on retry?idempotency_key, required on write and paid toolsSume docs

Why a stronger agent needs the cap more

An agent that completes more tasks also makes more calls without a person watching. In the Studio chat UI, an interactive spend-approval prompt protects the wallet. A backend caller does not get that prompt, and Sume's docs say the cap replaces it. Set it per run to the most you accept for that one task.

Batched tool calls make the same point from the other side. If the model fires five paid calls in one turn, each call is its own paid submit with its own idempotency_key. The key dedupes retries. It does not limit total spend.

A start call with the cap in the body

This helper refuses to build a request without a positive cap. It then posts to the real endpoint with your key from the environment. It uses only the standard library.

import json, os, urllib.request

def start_run(instruction: str, cap_usd: float) -> dict:
    if cap_usd <= 0:
        raise ValueError("set a positive generation_spend_cap_usd")
    body = {"instruction": instruction,
            "generation_spend_cap_usd": cap_usd}
    req = urllib.request.Request(
        "https://api.sume.com/v1/agent/completions",
        data=json.dumps(body).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json"},
        method="POST")
    with urllib.request.urlopen(req) as resp:
        return json.load(resp)["data"]

if __name__ == "__main__":
    run = start_run("Make one 6 second product teaser", 3)
    print(run["id"], run["status"], run["usage"])

Reading the receipt after the run

The 202 receipt already carries usage.generation_spend_cap_usd_micros, so the cap you set is echoed back before any work starts. After the run reaches a terminal status, usage records what was spent. Compare the two numbers for each task type and tighten caps that you never come close to using.

A run that returns usage: null means Sume could not read the spend, which is different from zero. Treat null as unknown and check GET /v1/usage before you assume the run was free. This habit matters more with a capable model, because capable agents finish more of the work you hand them, and a loose cap is only discovered when the invoice arrives.

What to do

Treat the vendor score as a reason to give the agent harder work, not larger budgets.

  • Choose the cap per task from the metered rates on the API pricing page, as the Sume docs advise.
  • Send an Idempotency-Key on the create call so a retried tool call returns the original receipt.
  • Poll the status_url from the 202 receipt, or register a webhook, instead of re-submitting.
  • Read usage on the receipt after the run to compare the cap with what was spent.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume