Send more than 100 spreadsheet rows to a Sume Format in chunks

One Sume bulk queue takes 1 to 100 items. For a 450-row sheet, split it into chunks and key each chunk so a retry never doubles spend.

4 min readSume
All posts

What do I do when my sheet has more than 100 rows?

POST /v1/formats/{handle}/{slug}/bulk-runs accepts items with 1 to 100 entries and a concurrency from 1 to 16. A 450-row sheet is therefore five queues: four of 100 rows and one of 50. Each item has the same body as a single run, so every row carries its own input, and optionally its own instruction or communication.webhook_url.

The queues are independent. Sume has no parent object tying them together, so your code owns the batch: the split, the keys, and the bookkeeping of which chunk got which queue id.

How do I make each chunk safe to retry?

Send an Idempotency-Key header per queue. A replay with the same payload returns 202 and the original queue instead of starting a second one. A different payload under the same key is a 409. The queue's key is scoped to your account and that Format, and the payload hash covers concurrency and items.

So the key should name the chunk and the sheet it came from. Hash the chunk's items into the key and an edited chunk gets a new key, while an unchanged chunk replays safely. If you would rather get the 409 as a guard against edits, leave the hash out and key on the batch id and chunk number alone.

import hashlib, json, os, requests

URL = "https://api.sume.com/v1/formats/acme/weekly-promo/bulk-runs"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def chunks(rows, size=100):
    for i in range(0, len(rows), size):
        yield i // size, rows[i:i + size]

def submit(batch_id, rows, concurrency=4):
    queues = {}
    for n, part in chunks(rows):
        items = [{"input": r} for r in part]
        digest = hashlib.sha256(
            json.dumps(items, sort_keys=True).encode()).hexdigest()[:12]
        headers = {**H, "Idempotency-Key": f"{batch_id}-c{n}-{digest}"}
        body = {"concurrency": concurrency, "items": items}
        r = requests.post(URL, json=body, headers=headers, timeout=30)
        r.raise_for_status()
        queues[n] = r.json()["data"]["id"]
    return queues

How should I set concurrency across chunks?

Run chunks one after another or in parallel, but remember the ceiling is your plan's generation concurrency: 1 on Free, 4 on Pro, 8 on Startup and 20 on Scale. Five queues each asking for 16 on a Pro workspace do not run 80 jobs at once; extra work waits for a slot.

A modest concurrency per queue, with queues submitted sequentially, keeps the order of completion predictable.

  • Store the returned queue id next to the chunk number the moment the 202 arrives.
  • Check counts.failed on each queue, because completed only means every item is terminal.
  • Record the spreadsheet row range per chunk so item index maps back to a row.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume