Bash: split a TSV into 100-row Sume bulk queues with jq and curl

A short shell script that cuts a product TSV into 100-row chunks, builds each bulk body with jq, and posts it under a chunk-named idempotency key.

5 min readSume
All posts

Export the sheet as tab-separated text, drop the header, cut it with split -l 100, and turn each piece into a bulk body with jq. A Sume bulk queue takes 1 to 100 items, so one file of 100 lines is one POST /v1/formats/{handle}/{slug}/bulk-runs. Name the idempotency key after the batch label and the chunk file, and a rerun of the script replays instead of rendering twice.

Tabs are the choice here because jq has no CSV parser. Splitting on commas breaks the first time a product title contains one. A tab inside a cell is rare, and the export step can strip it.

The script

products.tsv has a header row, then sku<TAB>product_url per line. The split output files are chunk_aa, chunk_ab and so on, which sort in order and make stable key suffixes.

set -euo pipefail
FORMAT=acme/product-promo; BATCH=bf-2026-v1
rm -f chunk_*
tail -n +2 products.tsv | split -l 100 - chunk_
for f in chunk_*; do
  jq -R -s -c '{concurrency: 4, items: (split("\n") | map(select(length > 0) | split("\t")
      | {instruction: "Make a 9:16 product ad.",
         input: {sku: .[0], product_url: .[1]},
         generation_spend_cap_usd: 5}))}' "$f" |
  curl -sS --fail-with-body -X POST \
    "https://api.sume.com/v1/formats/$FORMAT/bulk-runs" \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: $BATCH-$f" -d @- |
  jq -c --arg f "$f" '{file: $f, queue: .data.id, total: .data.counts.total}'
done > queues.jsonl

Reading the output

queues.jsonl gets one line per chunk. total should equal the number of lines in that chunk file, which is a cheap check that no row was dropped. Keep the file: the API has no endpoint that lists queues, so this is the record of which queue holds which rows. Item index in each queue matches the line order in its chunk file.

--fail-with-body makes curl exit non-zero on a 4xx or 5xx and still print the error envelope, so set -e stops the loop at the first bad chunk. That is what you want for a 400 invalid_request, because the response carries details.index naming the bad item. Nothing was created for that chunk and nothing was charged. Fix the row and run the script again: the chunks that already succeeded replay with the same keys.

Responses to expect from one chunk (read 2026-10-07)
StatusCodeWhat to do
202(queue receipt)Record data.id; the first concurrency items are already running
400invalid_requestAn item names none of instruction, input, previous_run_id, attachments; read details.index
409idempotency_conflictThe chunk changed since the first send; details.queue_id names the original
429rate_limitedWait retry-after seconds, then resend with the same key
402insufficient_creditsFund the wallet; the key is released, so the same key works afterwards

Keys, and what shifts a chunk

The key bf-2026-v1-chunk_aa is tied to the content of the first 100 lines. If a product is added near the top of the TSV, every later chunk shifts by one row, the bodies differ, and the same key now returns 409 idempotency_conflict with details.queue_id pointing at the old queue. That is the intended protection against sending two different jobs under one key. Change BATCH for a new export, and keep the old queues.jsonl for the old one.

A key can be up to 255 characters and belongs to one Format, so the same BATCH-chunk string sent to a second Format would start a second set of runs.

Limits that still apply

Each item's input can have at most 64 top-level keys and 2 MiB, and the whole request body is capped at 4 MiB. Two columns per row is nowhere near either. If a sheet has dozens of columns, nest them under one key as described in the 64-column post.

The rm -f chunk_* line matters: a leftover chunk_ac from a longer earlier export would be posted as if it were part of this batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume