Python preflight for a Sume bulk body: count, keys, bytes, caps
Check a Sume bulk-run body offline before the POST: 1 to 100 items, concurrency 1 to 16, 64 input keys, 2 MiB input, 4 MiB body and a spend cap in range.

A bulk create validates every item before it makes a queue, so one bad row out of 100 turns the whole call into a 400 invalid_request with the bad position in details.index. You can catch the same mistakes on your own machine in a short function: concurrency from 1 to 16, one to 100 items, each item carrying at least one of instruction, input, previous_run_id or attachments, input an object of at most 64 top-level keys and 2 MiB, a spend cap above 0 and at most 500, and a whole body under 4 MiB.
The rules below come from the Format API reference. The function does not replace the server's check, which also knows your Format's schema and your plan. It exists so a 3,000-row spreadsheet fails in your terminal in a second, with every problem listed, instead of one problem at a time through the API.
The checker
Pass a JSON file that holds the items array. The exit status is 1 when anything is wrong, so it also works as a CI step in front of the real call.
import json, sys
def size(x):
return len(json.dumps(x, separators=(",", ":"), ensure_ascii=False).encode())
def check(items, concurrency=4):
errs = []
if not 1 <= concurrency <= 16:
errs.append("concurrency must be 1 to 16")
if not 1 <= len(items) <= 100:
errs.append(f"items must be 1 to 100, got {len(items)}")
for i, it in enumerate(items):
if not any(k in it for k in ("instruction", "input", "previous_run_id", "attachments")):
errs.append(f"item {i}: no instruction, input, previous_run_id or attachments")
if "output_schema" in it and "response_format" in it:
errs.append(f"item {i}: output_schema and response_format together")
cap = it.get("generation_spend_cap_usd")
if cap is not None and not 0 < cap <= 500:
errs.append(f"item {i}: spend cap must be above 0 and at most 500")
inp = it.get("input")
if inp is not None and not isinstance(inp, dict):
errs.append(f"item {i}: input must be an object")
elif inp and (len(inp) > 64 or size(inp) > 2097152):
errs.append(f"item {i}: input over 64 keys or 2097152 bytes")
if size({"concurrency": concurrency, "items": items}) > 4 * 1024 * 1024:
errs.append("body over 4 MiB")
return errs
if __name__ == "__main__":
problems = check(json.load(open(sys.argv[1])))
print("\n".join(problems) or "ok")
sys.exit(1 if problems else 0)What each check mirrors
The input size is measured on the compact serialization: no spaces after commas or colons, and UTF-8 bytes rather than characters. That is how the docs state the 2097152-byte limit, and it is why size() passes separators and ensure_ascii=False. A default json.dumps pads with spaces and escapes non-ASCII characters, which makes a payload look larger than the server counts it.
The two body-level limits behave differently. The 2 MiB input limit is a 400. The 4 MiB request limit is a 413 payload_too_large with details.limit_bytes, and the fix is to send media by URL, or by a shared asset id, rather than inline.
The spend cap check accepts any value above 0 and up to 500. A cap of 0 is rejected, and null is treated as the maximum of 500. Leave the field off and the run inherits the Format's own cap instead.
| Rule | Limit | API answer when broken |
|---|---|---|
concurrency | 1 to 16 | 400 invalid_request |
items | 1 to 100 | 400 invalid_request |
| Item content | at least one of instruction, input, previous_run_id, attachments | 400 with details.index |
input shape | object, at most 64 top-level keys | 400 with details.index |
input size | 2097152 compact UTF-8 bytes | 400 with details.index |
generation_spend_cap_usd | above 0, at most 500 | 400 |
| Request body | 4 MiB | 413 payload_too_large |
What it cannot know
The checker does not look inside attachment URLs. One dead URL in an item still fails the whole create, so if your rows carry media links, fetch them first. It also cannot know whether the wallet can fund the batch: that is a 402 insufficient_credits at create, and nothing runs.
A clean result is permission to send, not a promise it will succeed. A queue that is accepted can still end with failed items, since each child is its own run.
Sources
Related posts
More in Developers
- queue_full 429 on a Sume submit: the reservation is released
A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
- Redact faces and license plates: Pillow first, AI edit only to replace
For redaction use Pillow boxes you control; use an AI mask edit on openai/gpt-image-2.5 only to replace a plate or face, from $0.0094 per image on Sume.
- Redeliver a missed video webhook after a bad deploy: one Sume call
Receiver down when the video finished? POST /v1/jobs/{job_id}/webhook/redeliver re-sends job.completed with a fresh signature. Scope, statuses, pitfalls.
Written by Sume