Sonnet 5.5 strict tool schema for a capped Sume Agent Completion

Sonnet 5.5 rejects forced tool_choice, so use auto plus strict: true. A tool schema that makes generation_spend_cap_usd mandatory for Sume Agent Completions.

6 min readSume
All posts

On Claude Sonnet 5.5 you cannot force a tool: tool_choice types any and tool return a 400, so you send auto and mark the tool strict: true. To let Claude start a Sume Agent Completion safely, define a client tool whose schema makes generation_spend_cap_usd required and limits it to a short enum, then have your handler call POST /v1/agent/completions and return the receipt. The cap stays a number you chose, not a number the model invented.

What Sonnet 5.5 changed

Anthropic's migration guide says Sonnet 5.5 rejects tool_choice of type any or tool with a 400, including on the token-counting endpoint. It tells you to send {"type": "auto"}, mark the tool strict: true, and say in the prompt when to call it, because the model may now answer without calling the tool.

Strict tool use constrains sampling so the tool input matches your JSON Schema; the page says the schema subset needs additionalProperties: false on every object, and its own examples use enum and required.

The tool definition

Sume's Agent Completions require exactly one of instruction or messages, and generation_spend_cap_usd has no default: omit it and the request fails with 400 invalid_request. A schema that mirrors that removes the failure class before it reaches Sume.

{
  "name": "start_sume_run",
  "description": "Start a Sume Agent Completion for one task. Call only when the user asked Sume to make or analyze media.",
  "strict": true,
  "input_schema": {
    "type": "object",
    "properties": {
      "instruction": {"type": "string"},
      "generation_spend_cap_usd": {"type": "number", "enum": [1, 2, 5]}
    },
    "required": ["instruction", "generation_spend_cap_usd"],
    "additionalProperties": false
  }
}

The handler is where the key lives

Your server receives the tool call, adds the secret and an Idempotency-Key, and posts to Sume. Replaying a key returns the original receipt with idempotency_hit: true, and reusing it with a different payload returns 409 idempotency_conflict, per the docs. Keep SUME_API_KEY in an environment variable on your server; the model never sees it. The key needs the agent_completions:write scope, and service-account keys cannot create completions.

The response is a 202 run receipt with status_url, not a chat answer, so return the run id and status URL as the tool result and let Claude tell the user the render is in progress.

Why an enum instead of a free number

A strict schema guarantees type-correct input, not sensible input. An enum of three ceilings means a model slip cannot become a 500 dollar cap. If you want the user to pick, show the enum values in the UI and pass the choice into the prompt.

Guarantees by layer (read 2026-10-04)
LayerGuaranteedNot guaranteed
strict: true on the toolInput matches the JSON SchemaThat calling the tool was the right choice
tool_choice: autoClaude may answer without the toolThat it calls the tool when you expect
Sume cap fieldRun spend ceiling you setAnything about the model's own token cost

Check before you ship

Create a fresh key in the dashboard if your current one predates Agent Completions: older keys fail with 403 insufficient_scope and scopes cannot be added to an existing key, as the docs describe. See Authentication and Safe automation for key handling and logging rules. Do not log signed URLs from the receipt.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume