Sizing generation_spend_cap_usd for an editing Agent Completion
generation_spend_cap_usd is required on every Agent Completion. Price the media jobs, then add headroom: edits of $0.22 to $1.24 suggest caps of $1 to $3.

generation_spend_cap_usd is required on every POST /v1/agent/completions call and has no default, so omitting it returns a 400. Size it by pricing the deterministic jobs you expect from the fixed rates, then add headroom for one retry of the costly step.
The rates below come from the Sume model pages read on 2026-10-03. The headroom multiplier is my judgment, not a documented figure, and the cap covers generation spend only; the Agent's own LLM turn is billed separately.
Price the plan first
Write down the jobs the instruction asks for and multiply by the fixed rate: trim is $0.02 per job, a video filter encode $0.02, captions $0.20 per job up to 60 seconds, timeline render $0.10 per output minute rounded up, compose $0.02 and audio detach $0.01. Speech-to-text is $0.01 per audio minute.
The Agent's reasoning also costs money, which is why the total debit will exceed the generation figure. Do not put video inspect on this list. It is metered on compute, so it has no fixed price to multiply; leave extra room if the task uses it.
| Instruction | Jobs | Expected | Cap I would send (my judgment) |
|---|---|---|---|
| Burn given text on one clip | 1 caption | $0.20 | $1 |
| Dim then caption one clip | 1 filter, 1 caption | $0.22 | $1 |
| Weekly recap from six clips | 6 trims, 1 render, 1 caption | $0.42 | $1 |
| Product teaser trio | 3 trims, 2 filters, 3 captions | $0.70 | $2 |
| Talk to four vertical shorts | 3 detaches, 25 min STT, 4 trims, 4 crops, 4 captions | $1.24 | $3 |
What the cap does and does not do
The cap is the ceiling on generation spend for that run; the receipt reports it as usage.generation_spend_cap_usd_micros so you can confirm what was applied. A cap that is too tight risks a partial result, because a single retried caption job alone is $0.20, and a cap sized to the bare plan leaves nothing for any retry at all. Check the Agent Completions page for how a run behaves when it reaches the cap before you rely on a specific outcome.
Too loose is the other failure. The Agent is free to choose tools, and a generous cap lets a vague instruction buy video generation you did not ask for. Say in the instruction which tools to use and that no new generation is wanted.
- State the allowed tools and the clip URLs in the instruction.
- Use
output_schemaso the result is machine-readable. - Set the cap per call; it is not remembered between calls.
Verify after the first run
Read GET /v1/usage?run_id= for the finished run and compare debited_usd_micros with your table. The receipt's billable_amount_usd_micros is generation only, so it will sit below the debit by the Agent's turns. If the generation figure is above your expected number, look at the operation breakdown in the summary, tighten the instruction, and lower the cap.
Sources
Related posts
More in Agents
- Sume Agent Completions request: required cap, model sume-agent
The minimum valid POST /v1/agent/completions body: instruction or messages, no assistant turns, model sume-agent, required generation_spend_cap_usd.
- Sume Agent Completions or a Format run: which should code call?
Pick between POST /v1/agent/completions and a Format run: open-ended instruction versus a reusable recipe, fresh thread each time, schema output, and cost cap.
- Sume output_schema name: slashes allowed, rewritten upstream
Names like sume/action-image-v1 are accepted (A-Z a-z 0-9 . _ / -, up to 64 chars) and echoed back unchanged. Only the structuring call sees a rewritten name.
- Sume schedule: cron or API call is fixed when you create it
A Scheduled agent's trigger_type cannot change after creation. Cron schedules can also accept API runs; API-only ones never gain a cadence. How to choose.
Written by Sume