Pandas DataFrame to a Sume bulk queue in 100-row chunks
Turn a product DataFrame into Sume Format bulk queues: one item per row, 100 rows per queue, a stable key per chunk, and SKU order saved beside each queue id.

To send a pandas DataFrame to Sume, convert each row to one item (instruction and input), slice the frame into 100-row pieces, and POST each piece to /v1/formats/{handle}/{slug}/bulk-runs with its own Idempotency-Key. A queue takes 1 to 100 items, so a 250-row sheet is three queues. Save the queue id next to the SKU list of that slice, because the queue reports items by position, not by SKU.
The snippet below does all of that with the standard library for HTTP and pandas only for the frame.
What one queue accepts
These are the limits that shape the slicing code. Each item is the same body as a single run, so the per-run input limits apply to every row.
| Field | Rule | Where it bites a DataFrame |
|---|---|---|
items | 1 to 100 entries, in order | Slice with iloc[start:start + 100] |
concurrency | Integer 1 to 16 | Required; there is no default |
input per item | JSON object, at most 64 top-level keys and 2 MiB | Do not pass a 70-column row as-is |
| Request body | 4 MiB maximum | Long instruction text times 100 adds up |
| Item content | Needs one of instruction, input, previous_run_id, attachments | An all-blank row makes the whole create fail with 400 |
The script
Set SUME_FORMAT to handle/slug, and SUME_API_KEY to a key that has formats:write. fillna("") matters: pandas holds a missing cell as a float that is not valid JSON, and an empty string is.
import json, os, urllib.request
import pandas as pd
URL = f"https://api.sume.com/v1/formats/{os.environ['SUME_FORMAT']}/bulk-runs"
BATCH = "bf-2026-v1" # bump only when you want a re-run
def row_item(row):
return {"instruction": "Make a 9:16 product ad from the page.",
"input": {"sku": row["sku"], "product_url": row["url"]},
"generation_spend_cap_usd": 5}
def create_queue(items, n):
req = urllib.request.Request(URL, method="POST",
data=json.dumps({"concurrency": 4, "items": items}).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": f"{BATCH}-chunk-{n}"})
with urllib.request.urlopen(req) as r:
return json.load(r)["data"]
def main():
df = pd.read_csv("products.csv").fillna("")
for n, start in enumerate(range(0, len(df), 100)):
part = df.iloc[start:start + 100]
q = create_queue([row_item(r) for r in part.to_dict("records")], n)
print(json.dumps({"queue": q["id"], "skus": list(part["sku"])}))
main()Why the SKU list goes beside the queue id
Each queue item has an index that matches its position in the items array you sent, and a run_id once it starts. The structured output of a run is built from what the run made, not from your input, so a SKU you sent will not come back in it unless the run repeats it. The printed line above is your join table: item 17 of that queue is the 18th SKU of that slice.
The cap in row_item is a per-item ceiling that you choose. Sume does not set a queue-level limit, so the worst case for a chunk is the sum of the item caps.
After the 202
A 202 means the queue exists and the first concurrency items are already running. The rest start as slots free up, so a 100-row chunk at concurrency: 4 keeps four runs in flight until the list drains. Workspace generation concurrency still applies to the children, so a larger window does not mean more parallel renders than your plan allows.
Poll GET /v1/format-run-queues/{id} (the status_url on the receipt) with formats:read. When status is completed, every item is terminal, which is not the same as every item succeeded. Look at counts.failed and counts.canceled, and read the child run receipt for the reason. The queue has no webhook, so progress at queue level is a poll; each item can still register its own communication.webhook_url.
Re-running a chunk
The key is bf-2026-v1-chunk-0, not a random UUID. If the script dies after the first POST and you start it again, the same key with the same body returns 202 and the queue that already exists, with no second charge. If you edited rows and send the same key, you get 409 idempotency_conflict, and details.queue_id names the original queue. Change BATCH when you really want new runs.
A bad row fails the whole create with 400 invalid_request and details.index, before any queue exists and before any spend. Fix that row and resend. For the failure cases after the 202, see what happens when the wallet runs dry mid-queue.
Sources
Related posts
More in Developers
- Pin the model id in an ad test: sume/auto follows the catalog
sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.
- Poll hundreds of AI jobs without a thundering herd: jitter and budgets
Poll many Sume jobs without synchronized bursts: jitter, next_poll_after_seconds, per-plan read budgets, and the math on how much polling a plan can absorb.
- Polling 200 video jobs every 30 s is 400 calls a minute: do this
A poll loop copied to Sume's 30 s example turns 200 open jobs into 400 status calls a minute. Use next_poll_after_seconds, backoff, or a webhook plus sweep.
- Portuguese speech to text API: Sume STT language_code pt or pt-BR
Transcribe Portuguese audio with Sume STT using language_code pt or pt-BR, then check the reported language and word times. $0.01 per audio minute.
Written by Sume