Stream a product feed NDJSON into 100-row Sume bulk bodies
Read an NDJSON product feed lazily with itertools.islice and write one bulk-run body per 100 rows, so a large file never has to fit in memory. Python stdlib.

A bulk queue takes at most 100 items, so a feed of any real size has to be cut into bodies of 100 rows. When the feed is NDJSON, one product per line, you can do it with a generator and itertools.islice without ever holding the whole file: take 100 lines, build the body, write it, and take the next 100. The last body is whatever is left over.
The script below writes bulk-000.json, bulk-001.json and so on. Each file is a complete create body, { "concurrency": 4, "items": [...] }, ready to POST to /v1/formats/{handle}/{slug}/bulk-runs.
The chunker
Each feed line needs sku, title and image. The input object is yours to design: the Format reads the keys it knows. Run it as python3 chunker.py feed.ndjson.
import itertools, json, sys
def bulk_bodies(lines, concurrency=4, size=100):
"""Yield one bulk-run body per `size` NDJSON lines without loading the file."""
rows = (json.loads(line) for line in lines if line.strip())
while batch := list(itertools.islice(rows, size)):
items = [{"input": {"sku": r["sku"], "title": r["title"], "image_url": r["image"]}}
for r in batch]
yield {"concurrency": concurrency, "items": items}
if __name__ == "__main__":
with open(sys.argv[1]) as feed:
for n, body in enumerate(bulk_bodies(feed)):
with open(f"bulk-{n:03}.json", "w") as out:
json.dump(body, out)
print(f"bulk-{n:03}.json", len(body["items"]))Notes on the shape
The walrus loop stops when islice returns an empty list, which is also what happens when the feed's line count is an exact multiple of 100: no empty body is produced. Blank lines are skipped before parsing.
Keep input small and flat. Each item allows 64 top-level keys and 2 MiB, and a body of 100 items must stay under 4 MiB, so the average item can be about 40 KiB. Product text and URLs are far below that. Put images in as URLs and not as inline data, and if every row uses the same packshot, upload it once and reference the asset id instead of repeating the URL.
Send each file with its own idempotency key, derived from the file name or from a hash of its contents, so a retried upload of bulk-001.json replays instead of rendering twice. The bodies are independent: each one is its own queue with its own concurrency window of 1 to 16.
| File | Items | Why |
|---|---|---|
bulk-000.json | 100 | First full body |
bulk-001.json | 100 | Second full body |
bulk-002.json | 5 | Remainder |
Before you send
A queue's worst-case spend is the sum of its items' spend caps, and an item without one inherits the Format's cap, so set generation_spend_cap_usd per item in the dict comprehension if the Format default is higher than a single product video should cost.
Validate the files before upload: a bad item gives a 400 with details.index and no queue exists.
Cutting by line count is the simplest rule and it is fine when rows are independent. If you want the files to be stable when the feed changes, sort by SKU first or cut by category, because adding one product at the top shifts every later boundary and changes every later file. Stable files matter if you key each upload on a hash of its contents.
Because the generator never holds more than 100 parsed rows, the same code handles a feed of ten thousand lines. The limit you will meet first is the plan's write rate, and a bulk create is one write however many items it carries, so 100 files of 100 rows is 100 writes.
Sources
Related posts
More in Developers
- SUME_API_AUTH_MODE: make the Sume CLI send Bearer instead of x-api-key
The Sume CLI sends x-api-key by default. Set SUME_API_AUTH_MODE=bearer when your client or network layer expects Authorization: Bearer. Never send both.
- A Sume job looks stuck: wait, cancel or poll the events?
Read status, then events. Queued and processing mean wait, cancel only works before generation starts, and a client timeout never cancels the job.
- sume login --no-browser on a remote server: approve the user_code URL
On an SSH box, sume login --no-browser prints the approval URL instead of opening a browser. After you approve the user_code, the CLI stores a CLI-scoped key.
- Sume media tools: which wait for a 200 and which always return 202
Video inspect defaults to sync; trim, filter, detach, compose and timeline to async; frames is always 202. The 30 s wait and where each result is read.
Written by Sume