SaaS AI video feature: spend cap per plan and per-customer keys

To embed Format runs in a SaaS plan, set generation_spend_cap_usd from the customer's tier and derive Idempotency-Key from customer, order and version.

6 min readSume
All posts

When you sell an AI video feature inside your own product, two request fields carry most of the safety: generation_spend_cap_usd, which you set from the customer's plan tier, and Idempotency-Key, which you derive from customer id, order id and a version you control. Together they stop a free-tier user from spending past their allowance and a double-click from starting two paid runs.

Both fields are described in Calling a Format and the Embed a Format cookbook. The cap values in the example are the cookbook's illustrations, not recommendations. A run that wants to spend past its cap fails with format_run_failed, so a tier whose runs keep failing is the first thing to check against usage.generation_spend_cap_usd_micros; size each cap above what one run of your Format really costs.

How should the spend cap follow your plans?

Every Format carries a generation spend cap, and a run can never spend past its own effective cap. The run request can name its own ceiling, up to the platform maximum of $500. The cookbook shows the cap as the natural place to express your own tiers: 0.5 for a free plan, 3 for pro, and nothing for enterprise so the run inherits the Format's own cap.

Read the Format's own cap from generation_spend_cap_usd_micros on GET /v1/formats/.... It is always a number, and a Format that never named one reports the platform default of $400. The effective cap for a given run comes back on the receipt as usage.generation_spend_cap_usd_micros.

What the cap field does (read 2026-10-03)
You sendThe run's cap
NothingThe Format's own cap
A number up to 500That number; above the Format's cap is honored, not clamped
nullThe platform maximum, $500; it lifts the ceiling, it does not remove it
0, or above 500400 invalid_request

What goes into the idempotency key?

The cookbook's rule is to hash stable identifiers from your own system: tenant id, order id, the Format slug, and a version you bump when you deliberately want a re-run. A uuid per request is called out as making the header decorative. A key built from the order id alone is also wrong, because two tenants with colliding order ids would share a run.

A replay with the same key and body returns 200 with the original receipt and idempotency_hit: true, with no second run and no second charge. The same key with a different body, even a different instruction, is 409 idempotency_conflict. Keys are scoped to one Format and may be up to 255 characters.

  • Store the returned run id against your own record before you answer the browser.
  • A create that failed with 402 or 503 releases the key, so retry with the same key after fixing the cause.
  • Two simultaneous requests with one key give 409 idempotency_key_in_use, which is retryable after about a second.

A server-side helper

This Python function derives the key, looks up the cap by plan, and starts the run. It uses a placeholder acme/product-promo Format, so substitute your own handle and slug, and keep the API key on the server.

import hashlib
import os
import requests

CAPS = {"free": 0.5, "pro": 3}  # illustrative; enterprise inherits the Format cap


def run_key(customer_id, order_id):
    raw = f"{customer_id}:{order_id}:product-promo:v1"
    return hashlib.sha256(raw.encode()).hexdigest()[:40]


def start_run(customer_id, order_id, plan, product_url):
    body = {
        "instruction": "Vertical 9:16 product promo. No captions.",
        "input": {"product_url": product_url},
    }
    cap = CAPS.get(plan)
    if cap is not None:
        body["generation_spend_cap_usd"] = cap
    r = requests.post(
        "https://api.sume.com/v1/formats/acme/product-promo/runs",
        headers={
            "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
            "Idempotency-Key": run_key(customer_id, order_id),
        },
        json=body,
        timeout=30,
    )
    r.raise_for_status()
    return r.json()["data"]["id"]

What the cap does not cover

Caps bound generation spend. The terminal receipt reports what the run spent against that ceiling as usage.billable_amount_usd_micros, which excludes the agent's own LLM turn. So it is not the run's total cost and not an invoice. Bill your customers from your own records and reconcile against GET /v1/usage.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume