Sizing generation_spend_cap_usd for an editing Agent Completion

generation_spend_cap_usd is required on every Agent Completion. Price the media jobs, then add headroom: edits of $0.22 to $1.24 suggest caps of $1 to $3.

4 min readSume
All posts

generation_spend_cap_usd is required on every POST /v1/agent/completions call and has no default, so omitting it returns a 400. Size it by pricing the deterministic jobs you expect from the fixed rates, then add headroom for one retry of the costly step.

The rates below come from the Sume model pages read on 2026-10-03. The headroom multiplier is my judgment, not a documented figure, and the cap covers generation spend only; the Agent's own LLM turn is billed separately.

Price the plan first

Write down the jobs the instruction asks for and multiply by the fixed rate: trim is $0.02 per job, a video filter encode $0.02, captions $0.20 per job up to 60 seconds, timeline render $0.10 per output minute rounded up, compose $0.02 and audio detach $0.01. Speech-to-text is $0.01 per audio minute.

The Agent's reasoning also costs money, which is why the total debit will exceed the generation figure. Do not put video inspect on this list. It is metered on compute, so it has no fixed price to multiply; leave extra room if the task uses it.

Expected media cost for five editing instructions and a suggested cap (read 2026-10-03)
InstructionJobsExpectedCap I would send (my judgment)
Burn given text on one clip1 caption$0.20$1
Dim then caption one clip1 filter, 1 caption$0.22$1
Weekly recap from six clips6 trims, 1 render, 1 caption$0.42$1
Product teaser trio3 trims, 2 filters, 3 captions$0.70$2
Talk to four vertical shorts3 detaches, 25 min STT, 4 trims, 4 crops, 4 captions$1.24$3

What the cap does and does not do

The cap is the ceiling on generation spend for that run; the receipt reports it as usage.generation_spend_cap_usd_micros so you can confirm what was applied. A cap that is too tight risks a partial result, because a single retried caption job alone is $0.20, and a cap sized to the bare plan leaves nothing for any retry at all. Check the Agent Completions page for how a run behaves when it reaches the cap before you rely on a specific outcome.

Too loose is the other failure. The Agent is free to choose tools, and a generous cap lets a vague instruction buy video generation you did not ask for. Say in the instruction which tools to use and that no new generation is wanted.

  • State the allowed tools and the clip URLs in the instruction.
  • Use output_schema so the result is machine-readable.
  • Set the cap per call; it is not remembered between calls.

Verify after the first run

Read GET /v1/usage?run_id= for the finished run and compare debited_usd_micros with your table. The receipt's billable_amount_usd_micros is generation only, so it will sit below the debit by the Agent's turns. If the generation figure is above your expected number, look at the operation breakdown in the summary, tighten the instruction, and lower the cap.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume