SaaS AI video feature: spend cap per plan and per-customer keys
To embed Format runs in a SaaS plan, set generation_spend_cap_usd from the customer's tier and derive Idempotency-Key from customer, order and version.

When you sell an AI video feature inside your own product, two request fields carry most of the safety: generation_spend_cap_usd, which you set from the customer's plan tier, and Idempotency-Key, which you derive from customer id, order id and a version you control. Together they stop a free-tier user from spending past their allowance and a double-click from starting two paid runs.
Both fields are described in Calling a Format and the Embed a Format cookbook. The cap values in the example are the cookbook's illustrations, not recommendations. A run that wants to spend past its cap fails with format_run_failed, so a tier whose runs keep failing is the first thing to check against usage.generation_spend_cap_usd_micros; size each cap above what one run of your Format really costs.
How should the spend cap follow your plans?
Every Format carries a generation spend cap, and a run can never spend past its own effective cap. The run request can name its own ceiling, up to the platform maximum of $500. The cookbook shows the cap as the natural place to express your own tiers: 0.5 for a free plan, 3 for pro, and nothing for enterprise so the run inherits the Format's own cap.
Read the Format's own cap from generation_spend_cap_usd_micros on GET /v1/formats/.... It is always a number, and a Format that never named one reports the platform default of $400. The effective cap for a given run comes back on the receipt as usage.generation_spend_cap_usd_micros.
| You send | The run's cap |
|---|---|
| Nothing | The Format's own cap |
| A number up to 500 | That number; above the Format's cap is honored, not clamped |
| null | The platform maximum, $500; it lifts the ceiling, it does not remove it |
| 0, or above 500 | 400 invalid_request |
What goes into the idempotency key?
The cookbook's rule is to hash stable identifiers from your own system: tenant id, order id, the Format slug, and a version you bump when you deliberately want a re-run. A uuid per request is called out as making the header decorative. A key built from the order id alone is also wrong, because two tenants with colliding order ids would share a run.
A replay with the same key and body returns 200 with the original receipt and idempotency_hit: true, with no second run and no second charge. The same key with a different body, even a different instruction, is 409 idempotency_conflict. Keys are scoped to one Format and may be up to 255 characters.
- Store the returned run id against your own record before you answer the browser.
- A create that failed with
402or503releases the key, so retry with the same key after fixing the cause. - Two simultaneous requests with one key give
409 idempotency_key_in_use, which is retryable after about a second.
A server-side helper
This Python function derives the key, looks up the cap by plan, and starts the run. It uses a placeholder acme/product-promo Format, so substitute your own handle and slug, and keep the API key on the server.
import hashlib
import os
import requests
CAPS = {"free": 0.5, "pro": 3} # illustrative; enterprise inherits the Format cap
def run_key(customer_id, order_id):
raw = f"{customer_id}:{order_id}:product-promo:v1"
return hashlib.sha256(raw.encode()).hexdigest()[:40]
def start_run(customer_id, order_id, plan, product_url):
body = {
"instruction": "Vertical 9:16 product promo. No captions.",
"input": {"product_url": product_url},
}
cap = CAPS.get(plan)
if cap is not None:
body["generation_spend_cap_usd"] = cap
r = requests.post(
"https://api.sume.com/v1/formats/acme/product-promo/runs",
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": run_key(customer_id, order_id),
},
json=body,
timeout=30,
)
r.raise_for_status()
return r.json()["data"]["id"]
What the cap does not cover
Caps bound generation spend. The terminal receipt reports what the run spent against that ceiling as usage.billable_amount_usd_micros, which excludes the agent's own LLM turn. So it is not the run's total cost and not an invoice. Bill your customers from your own records and reconcile against GET /v1/usage.
Sources
Related posts
More in Formats
- Season output schema: episode videos and final cut as SumeMediaFile
Bind an output_schema with SumeMediaFile fields so a Sume run returns typed episode videos. The URL gate, 10% duration check and parts-versus-cut rule.
- Series bible for a Format: SKILL.md index, references and run input
Where to keep a series bible in a Sume Format: a short SKILL.md index, detail in references/*, and only the per-episode beat in the run input.
- Slideshow Format: a holiday gift guide from up to 30 product images
Send up to 30 product images to Sume's slideshow Format for a gift-guide clip, then check Pinterest's video ad specs before you promote it.
- Bulk Format runs: 100 items, 16 at once, what completed means
Sume bulk runs take 1 to 100 items at concurrency 1 to 16. A queue marked completed means every item is terminal, not that every item succeeded.
Written by Sume