Cap what one Format run can spend with generation_spend_cap_usd

Every Sume Format run has a generation spend cap. Send generation_spend_cap_usd to set it: up to 500, null runs at 500, 0 is rejected. What happens at the cap.

4 min readSume
All posts

To cap what one Format run can spend, send generation_spend_cap_usd in the body of POST /v1/formats/{handle}/{slug}/runs. Omit it to inherit the Format's own cap; send a number up to 500 for this run's ceiling; null runs at the $500 platform maximum; 0 or anything above 500 is 400. A run can never spend past its effective cap.

Rules are from Create a run, read 2026-09-29.

What does it look like?

The receipt echoes the effective cap as usage.generation_spend_cap_usd_micros.

curl -X POST https://api.sume.com/v1/formats/acme/promo/runs \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: promo-cap-001" \
  -d '{
    "instruction": "One 9:16 clip, three shots",
    "input": { "product_name": "Aurora Headphones" },
    "generation_spend_cap_usd": 8
  }'

What does each value do?

Cap values, from Create a run, read 2026-09-29.
You sendThe run's cap
NothingThe Format's cap ($400 when it never set one)
A number up to 500That number; above the Format's own cap is honored, not clamped
null$500, the platform maximum
0, or above 500400

What happens at the cap?

The run ends failed and usage shows how close it got. The generic format_run_failed error is also what a run that wanted to spend past its cap lands on, so compare usage.billable_amount_usd_micros with usage.generation_spend_cap_usd_micros before you raise the brief.

What does the spend figure include?

Metered generation: video, image, avatar, voice and timeline work. It excludes the agent's own LLM turn, so it is not the run's total cost, and it is a receipt figure, not an invoice; GET /v1/usage and GET /v1/balance are the billing records. The docs suggest caps around $120 for production long-form runs and a few dollars for a single-scene retry, which are their examples, not a recommendation for you.

Do scheduled runs work the same way?

Not quite: a schedule has its own cap and a per-run override can only lower it. See the schedule cap.

How do I cap a whole bulk batch?

Each bulk item is the same body as a single run, so each item can carry its own generation_spend_cap_usd. A bulk request queues 1 to 100 items with a concurrency window of 1 to 16; poll GET /v1/format-run-queues/{id}. The queue has no webhook of its own, and completed means every item is terminal, not that every item succeeded, so check counts.failed. See bulk runs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume