Pandas DataFrame to a Sume bulk queue in 100-row chunks

Turn a product DataFrame into Sume Format bulk queues: one item per row, 100 rows per queue, a stable key per chunk, and SKU order saved beside each queue id.

5 min readSume
All posts

To send a pandas DataFrame to Sume, convert each row to one item (instruction and input), slice the frame into 100-row pieces, and POST each piece to /v1/formats/{handle}/{slug}/bulk-runs with its own Idempotency-Key. A queue takes 1 to 100 items, so a 250-row sheet is three queues. Save the queue id next to the SKU list of that slice, because the queue reports items by position, not by SKU.

The snippet below does all of that with the standard library for HTTP and pandas only for the frame.

What one queue accepts

These are the limits that shape the slicing code. Each item is the same body as a single run, so the per-run input limits apply to every row.

Bulk queue and per-item limits (read 2026-10-07)
FieldRuleWhere it bites a DataFrame
items1 to 100 entries, in orderSlice with iloc[start:start + 100]
concurrencyInteger 1 to 16Required; there is no default
input per itemJSON object, at most 64 top-level keys and 2 MiBDo not pass a 70-column row as-is
Request body4 MiB maximumLong instruction text times 100 adds up
Item contentNeeds one of instruction, input, previous_run_id, attachmentsAn all-blank row makes the whole create fail with 400

The script

Set SUME_FORMAT to handle/slug, and SUME_API_KEY to a key that has formats:write. fillna("") matters: pandas holds a missing cell as a float that is not valid JSON, and an empty string is.

import json, os, urllib.request
import pandas as pd

URL = f"https://api.sume.com/v1/formats/{os.environ['SUME_FORMAT']}/bulk-runs"
BATCH = "bf-2026-v1"  # bump only when you want a re-run

def row_item(row):
    return {"instruction": "Make a 9:16 product ad from the page.",
            "input": {"sku": row["sku"], "product_url": row["url"]},
            "generation_spend_cap_usd": 5}

def create_queue(items, n):
    req = urllib.request.Request(URL, method="POST",
        data=json.dumps({"concurrency": 4, "items": items}).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json",
                 "Idempotency-Key": f"{BATCH}-chunk-{n}"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)["data"]

def main():
    df = pd.read_csv("products.csv").fillna("")
    for n, start in enumerate(range(0, len(df), 100)):
        part = df.iloc[start:start + 100]
        q = create_queue([row_item(r) for r in part.to_dict("records")], n)
        print(json.dumps({"queue": q["id"], "skus": list(part["sku"])}))

main()

Why the SKU list goes beside the queue id

Each queue item has an index that matches its position in the items array you sent, and a run_id once it starts. The structured output of a run is built from what the run made, not from your input, so a SKU you sent will not come back in it unless the run repeats it. The printed line above is your join table: item 17 of that queue is the 18th SKU of that slice.

The cap in row_item is a per-item ceiling that you choose. Sume does not set a queue-level limit, so the worst case for a chunk is the sum of the item caps.

After the 202

A 202 means the queue exists and the first concurrency items are already running. The rest start as slots free up, so a 100-row chunk at concurrency: 4 keeps four runs in flight until the list drains. Workspace generation concurrency still applies to the children, so a larger window does not mean more parallel renders than your plan allows.

Poll GET /v1/format-run-queues/{id} (the status_url on the receipt) with formats:read. When status is completed, every item is terminal, which is not the same as every item succeeded. Look at counts.failed and counts.canceled, and read the child run receipt for the reason. The queue has no webhook, so progress at queue level is a poll; each item can still register its own communication.webhook_url.

Re-running a chunk

The key is bf-2026-v1-chunk-0, not a random UUID. If the script dies after the first POST and you start it again, the same key with the same body returns 202 and the queue that already exists, with no second charge. If you edited rows and send the same key, you get 409 idempotency_conflict, and details.queue_id names the original queue. Change BATCH when you really want new runs.

A bad row fails the whole create with 400 invalid_request and details.index, before any queue exists and before any spend. Fix that row and resend. For the failure cases after the 202, see what happens when the wallet runs dry mid-queue.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume