Sizing generation_spend_cap_usd for a tool call from GPT-6.1 Sol
generation_spend_cap_usd has no default on Sume Agent Completions. Size it per run from metered API pricing and clamp it in your GPT-6.1 Sol tool handler.

Size generation_spend_cap_usd as the most you would accept losing on one run, then check it against the metered rates for the media the run will create. There is no default: omit the field and the request fails with 400 invalid_request. The cap replaces the interactive spend approval that the chat UI shows, so when a model such as GPT-6.1 Sol chooses the number through a function tool, your handler should clamp it before it reaches Sume.
Why the cap exists
The Agent Completions page explains that an Agent Completion is an unattended agent with tools and access to your generation wallet. The cap is the substitute for the approval prompt a backend caller never sees, and the docs say to set it per run to the most you will spend. The receipt echoes it as usage.generation_spend_cap_usd_micros, so you can log the value that was actually applied. To size it, the docs point to the metered rates on the API pricing page.
A sizing routine
- List the assets you expect: for example four stills, one voiceover, one clip.
- Price each at the metered rate on the API pricing page, and add them.
- Add headroom for a retry or two, but not an open-ended amount.
- Set a hard ceiling in code, and use
min(requested, ceiling)for what you send. - Log the cap with the run id so a morning audit can compare cap and recorded spend.
Where the numbers live
Do not copy rates into a prompt or a blog table. Rates change, and a copy goes stale quietly. Read them from the pricing page when you size, and from GET /v1/usage afterwards, which the docs list as the usage ledger, the billing record of what you actually spent.
| Limit | Set by | Enforced by |
|---|---|---|
generation_spend_cap_usd | Your server, per run | Sume, on the run's generation spend |
| Ceiling constant in handler | Your code | Your code, before the request |
| Wallet balance | Your account | Admission: 402 insufficient_credits |
| Plan concurrency | Your plan | Admission: 429 queue_full or rate_limited |
What the cap does not cover
The docs describe the cap as a ceiling on generation spend for the run. They do not say it covers anything your own code pays elsewhere, such as the tokens for the model that calls your function tool, so budget that separately. For the admission errors a run can still hit, read Generation admission, and keep key handling in line with Safe automation.
Sources
Related posts
More in Pricing
- AI music cost per finished minute, compared
Lyria 3.5, ElevenLabs Music v2.5 and Suno Pro priced per minute or song, against Sume's flat $0.125 per accepted music generation.
- AI music for a 15-second bumper ad: per-minute vs flat pricing
A 15-second bumper needs 15 seconds of music. ElevenLabs at $0.15 a minute is about $0.04; Google lists $0.08 a song; Sume is a flat $0.125. When flat loses.
- What a three-minute AI song costs: ElevenLabs, Lyria and Sume
At list prices a three-minute track is $0.45 on ElevenLabs Music API ($0.15 per minute), $0.08 on Lyria 3.5 and $0.125 on Sume's Music Router.
- AI video price per frame at 24 fps: Wan, Omni, H3 Max, Veo Lite
A 24 fps second is 24 frames. Price per frame: Veo 3.1 Lite $0.0021, Sume Wan 3.0 480p $0.0026, H3 Max 768p $0.0042, Wan and Omni 720p $0.0052.
Written by Sume