Bulk queue spend ceiling: 100 items at $2 is $200; 16 live is $32

With generation_spend_cap_usd of $2 per bulk item, 100 items cap at $200 in total and 16 running items at $32 at once. Caps are ceilings, not prices.

5 min readSume
All posts

Give each bulk item a generation_spend_cap_usd of $2 and the queue cannot spend more than 100 x $2 = $200 in total, and at most 16 x $2 = $32 can be in flight at any moment at concurrency 16. A cap is a ceiling for one run, not a price, so the real bill is what the runs actually generate. The numbers help you decide the cap, not predict the invoice.

What the docs say about caps

Every Format has a generation spend cap, and a run can never spend more than its own effective cap. The item body is the same as a single-run body, so each item can name its own cap.

generation_spend_cap_usd on a run (read 2026-10-08)
You sendThe run's cap
nothingthe Format's cap (the platform default is $400 if the Format names none)
a number up to 500that number
nullthe platform maximum, $500
0, or above 500400 error

Two ceilings for one queue

The first is the total ceiling: items x cap. The second is the exposure at one moment: concurrency x cap. The second one is what you watch while the queue runs, because a failing batch burns through the in-flight runs first.

Ceilings for a $2 item cap
ItemsConcurrencyTotal ceilingIn-flight ceiling
10016100 x 2 = $20016 x 2 = $32
1004100 x 2 = $2004 x 2 = $8
20220 x 2 = $402 x 2 = $4

What a cap does when it is hit

The docs say that a run can never spend more than its own effective cap, and that a cap of zero is rejected because a run that cannot spend cannot deliver. So a cap is a ceiling for that one run. A child that does not finish is recorded on its own item, and the queue goes on with the rest. Read the run receipt (GET /v1/format-runs/{run_id}) and not just the queue row to see why.

This is the reason to set caps per item and not rely on the Format default. The default is the Format's own cap, or $400 if the Format names none. A queue of 100 under a $400 default has a total ceiling of $40,000, which is probably not the number that you meant.

How to choose the cap

Run one item alone first and read usage.billable_amount_usd_micros from its receipt. That is the actual spend of the run. Set the item cap at a small multiple of it, for example twice the cost. A cap that is too tight fails runs that were healthy; a cap near $500 offers little protection.

A narrower window also limits the damage of a bad prompt. If the first four items all fail, a window of 4 has spent four runs, while a window of 16 has spent sixteen. When you are not sure about the batch, start at 4 and raise the window in a later queue.

Build the body with the cap on every item

The script builds a queue body and prints both ceilings. It refuses a cap that the API would reject.

import json

def queue_body(rows, concurrency, cap_usd):
    if not 1 <= concurrency <= 16 or not 1 <= len(rows) <= 100:
        raise ValueError("concurrency 1-16 and 1-100 items")
    if not 0 < cap_usd <= 500:
        raise ValueError("cap must be above 0 and at most 500")
    items = [{"instruction": r, "generation_spend_cap_usd": cap_usd} for r in rows]
    print("total ceiling: $%.2f" % (len(items) * cap_usd))
    print("in-flight ceiling: $%.2f" % (concurrency * cap_usd))
    return {"concurrency": concurrency, "items": items}

body = queue_body([f"clip {i}" for i in range(1, 101)], 16, 2)
print(json.dumps(body["items"][0]))

Sources

Related posts

More in Formats

All Formats posts

Written by Sume