Can a per-run spend cap raise the limit? Formats yes, schedules no
Formats honor a per-run generation_spend_cap_usd above the Format cap, schedules clamp it, and Agent Completions require it. Defaults and edge cases.

It depends on the surface. On a Format run, a generation_spend_cap_usd above the Format's own cap is honored, not clamped. On a scheduled agent, a per-run value can only lower the schedule's cap, never raise it. On an Agent Completion there is no default: omit it and the request fails with 400 invalid_request.
How do the three surfaces differ?
All three take a per-run generation_spend_cap_usd, and the defaults and ceilings are different enough to cause surprises when you move code between them. The table is from the Sume docs for each surface.
| Surface | Default when omitted | Number above the stored cap | null | 0 |
|---|---|---|---|---|
| Format run | The Format's cap ($400 if it never set one) | Honored, up to $500 | Runs at the $500 platform maximum | 400 |
| Scheduled agent run | $1.00 per run | Clamped to the schedule's cap | Runs without the automation ceiling | Rejected |
| Agent Completion | None; request fails with 400 | Not applicable | Not applicable | Not applicable |
Why would a Format honor a higher cap?
The Format's cap is a default for runs that do not name one, and the docs say a run "can never spend past its own effective cap". A single-scene retry needs a fraction of a long-form run, and a rare big run may need more than the default, so the caller is allowed to set the ceiling for that run, up to $500.
The docs also warn that null "lifts the ceiling; it does not remove it": the run is capped at $500, not left uncapped. Production long-form runs are typically created with caps around $120.
Why do schedules clamp instead?
A schedule runs unattended on a clock. The schedule's cap is set by the person who made it, and the per-run override is clamped to min(request, schedule cap), so a stray API caller cannot raise it. The null case is narrower than it sounds: it removes the automation ceiling, but wallet balance, generation admission and org limits still apply.
curl -X POST "https://api.sume.com/v1/formats/$HANDLE/$SLUG/runs" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"instruction": "Make the 30-second spot", "generation_spend_cap_usd": 120}'How should I choose a cap?
Size it from the rate card on the API pricing page, plus headroom for retries, and set it per run rather than relying on a default. Read the effective cap back from the receipt as usage.generation_spend_cap_usd_micros. If you port a call from a Format to a schedule, expect the clamp: a $120 request against a $1.00 schedule cap runs at $1.00.
Sources
Related posts
More in Pricing
- Rask AI minute credits and 3x lip-sync vs a Sume dubbing pipeline
Rask bills 1 credit per video minute for Standard lip-sync, 3 for Enhanced. Sume has no dubbing endpoint; a chained pipeline runs about $0.55-$1.55 per 10 min.
- Realtime voice API cost per hour: Grok, GPT-Live-1, Gemini Live
Per-hour cost of realtime voice APIs from the vendors' own pages, set against Sume's async STT plus TTS jobs, with the arithmetic and its assumptions shown.
- Recraft API pricing: 1,000 API units per dollar, V4.1 Flash $0.007
Recraft sells prepaid API units at $1.00 for 1,000. V4.1 Flash is 7 units ($0.007) per raster image, V4.1 is 35 units ($0.035), V4.1 Pro is 210 units ($0.21).
- Recraft Studio plan includes all models: when per-image billing wins
Recraft says every model in Studio is included in its standard plans. That is a seat model; Sume bills each completed image to a wallet. Which suits you?
Written by Sume