Sizing generation_spend_cap_usd for a tool call from GPT-6.1 Sol

generation_spend_cap_usd has no default on Sume Agent Completions. Size it per run from metered API pricing and clamp it in your GPT-6.1 Sol tool handler.

5 min readSume
All posts

Size generation_spend_cap_usd as the most you would accept losing on one run, then check it against the metered rates for the media the run will create. There is no default: omit the field and the request fails with 400 invalid_request. The cap replaces the interactive spend approval that the chat UI shows, so when a model such as GPT-6.1 Sol chooses the number through a function tool, your handler should clamp it before it reaches Sume.

Why the cap exists

The Agent Completions page explains that an Agent Completion is an unattended agent with tools and access to your generation wallet. The cap is the substitute for the approval prompt a backend caller never sees, and the docs say to set it per run to the most you will spend. The receipt echoes it as usage.generation_spend_cap_usd_micros, so you can log the value that was actually applied. To size it, the docs point to the metered rates on the API pricing page.

A sizing routine

  • List the assets you expect: for example four stills, one voiceover, one clip.
  • Price each at the metered rate on the API pricing page, and add them.
  • Add headroom for a retry or two, but not an open-ended amount.
  • Set a hard ceiling in code, and use min(requested, ceiling) for what you send.
  • Log the cap with the run id so a morning audit can compare cap and recorded spend.

Where the numbers live

Do not copy rates into a prompt or a blog table. Rates change, and a copy goes stale quietly. Read them from the pricing page when you size, and from GET /v1/usage afterwards, which the docs list as the usage ledger, the billing record of what you actually spent.

Who sets which limit (read 2026-10-04)
LimitSet byEnforced by
generation_spend_cap_usdYour server, per runSume, on the run's generation spend
Ceiling constant in handlerYour codeYour code, before the request
Wallet balanceYour accountAdmission: 402 insufficient_credits
Plan concurrencyYour planAdmission: 429 queue_full or rate_limited

What the cap does not cover

The docs describe the cap as a ceiling on generation spend for the run. They do not say it covers anything your own code pays elsewhere, such as the tokens for the model that calls your function tool, so budget that separately. For the admission errors a run can still hit, read Generation admission, and keep key handling in line with Safe automation.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume