Stream a product feed NDJSON into 100-row Sume bulk bodies

Read an NDJSON product feed lazily with itertools.islice and write one bulk-run body per 100 rows, so a large file never has to fit in memory. Python stdlib.

5 min readSume
All posts

A bulk queue takes at most 100 items, so a feed of any real size has to be cut into bodies of 100 rows. When the feed is NDJSON, one product per line, you can do it with a generator and itertools.islice without ever holding the whole file: take 100 lines, build the body, write it, and take the next 100. The last body is whatever is left over.

The script below writes bulk-000.json, bulk-001.json and so on. Each file is a complete create body, { "concurrency": 4, "items": [...] }, ready to POST to /v1/formats/{handle}/{slug}/bulk-runs.

The chunker

Each feed line needs sku, title and image. The input object is yours to design: the Format reads the keys it knows. Run it as python3 chunker.py feed.ndjson.

import itertools, json, sys

def bulk_bodies(lines, concurrency=4, size=100):
    """Yield one bulk-run body per `size` NDJSON lines without loading the file."""
    rows = (json.loads(line) for line in lines if line.strip())
    while batch := list(itertools.islice(rows, size)):
        items = [{"input": {"sku": r["sku"], "title": r["title"], "image_url": r["image"]}}
                 for r in batch]
        yield {"concurrency": concurrency, "items": items}

if __name__ == "__main__":
    with open(sys.argv[1]) as feed:
        for n, body in enumerate(bulk_bodies(feed)):
            with open(f"bulk-{n:03}.json", "w") as out:
                json.dump(body, out)
            print(f"bulk-{n:03}.json", len(body["items"]))

Notes on the shape

The walrus loop stops when islice returns an empty list, which is also what happens when the feed's line count is an exact multiple of 100: no empty body is produced. Blank lines are skipped before parsing.

Keep input small and flat. Each item allows 64 top-level keys and 2 MiB, and a body of 100 items must stay under 4 MiB, so the average item can be about 40 KiB. Product text and URLs are far below that. Put images in as URLs and not as inline data, and if every row uses the same packshot, upload it once and reference the asset id instead of repeating the URL.

Send each file with its own idempotency key, derived from the file name or from a hash of its contents, so a retried upload of bulk-001.json replays instead of rendering twice. The bodies are independent: each one is its own queue with its own concurrency window of 1 to 16.

What the 205-row demo feed produced (read 2026-10-07)
FileItemsWhy
bulk-000.json100First full body
bulk-001.json100Second full body
bulk-002.json5Remainder

Before you send

A queue's worst-case spend is the sum of its items' spend caps, and an item without one inherits the Format's cap, so set generation_spend_cap_usd per item in the dict comprehension if the Format default is higher than a single product video should cost.

Validate the files before upload: a bad item gives a 400 with details.index and no queue exists.

Cutting by line count is the simplest rule and it is fine when rows are independent. If you want the files to be stable when the feed changes, sort by SKU first or cut by category, because adding one product at the top shifts every later boundary and changes every later file. Stable files matter if you key each upload on a hash of its contents.

Because the generator never holds more than 100 parsed rows, the same code handles a feed of ten thousand lines. The limit you will meet first is the plan's write rate, and a bulk create is one write however many items it carries, so 100 files of 100 rows is 100 writes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume