OpenAI batch output order differs from input: how Sume bulk keys rows

OpenAI batch output may not follow input order, so you join on custom_id. A Sume bulk queue lists items by index in submitted order. Here is the join for each.

5 min readSume
All posts

Do OpenAI batch results come back in the order you sent them? No. The OpenAI Batch guide (read 2026-10-04) says the lines in the output file may not match the order of the input file, so you map each result to its request with custom_id. A Sume bulk queue works the other way round: the queue receipt lists one item per submitted row, in the same order, each with a zero-based index. Your join key on Sume is the position you sent, not a label you invented.

That difference matters the week a retailer loads a spreadsheet of holiday SKUs and needs each finished video back on the right row. Get the join wrong and a product page shows another product's clip.

The fix is cheap in both systems, but it is a different fix. With OpenAI you carry an identifier through the request. With Sume you carry nothing and trust the order of the receipt, so the one thing you must never do is sort, filter or de-duplicate your rows between building the request and reading the queue.

What each API gives you to join on

The table compares the two join keys as documented. The Sume row comes from the Bulk runs page.

Result join keys, vendor page and Sume docs read 2026-10-04
OpenAI Batch APISume bulk run
Request label you choosecustom_id on every input lineNone. Each item is the body of one Format run
Order of resultsMay not match input orderSame order as submitted items
How to joinMatch output line to input by custom_iditems[i] belongs to the i-th row you sent
FailuresRecorded separately from successesItem status failed, with error and run_id (null if the run never started)
Max size per submission50,000 requests, 200 MB input file1 to 100 items

Join a Sume queue back to your rows

Because the queue is ordered, the safest pattern is to keep your rows in a list and zip them with items from the queue. Each item carries index, status, run_id and error. Do not rely on the order in which runs finish; the window starts the next queued item as soon as a slot opens, so completion order is whatever the work dictates. Position in items is stable.

If you want a label anyway, put it in the item's input object, since you choose that shape. The Format reads the keys it knows and ignores the rest, so a sku key costs nothing and makes your own logs readable. That is a convention you add, not a field Sume echoes back as a join key.

A runnable zip

This snippet takes the rows you submitted and a queue receipt you fetched from GET /v1/format-run-queues/{id}, and prints the status per row. It uses only the standard library.

import json

rows = [{"sku": "mug-01"}, {"sku": "mug-02"}, {"sku": "mug-03"}]
queue = json.loads('''{"items": [
  {"index": 0, "status": "completed", "run_id": "r0", "error": null},
  {"index": 1, "status": "failed", "run_id": null,
   "error": {"code": "format_run_failed_to_start", "message": "x"}},
  {"index": 2, "status": "running", "run_id": "r2", "error": null}]}''')

items = sorted(queue["items"], key=lambda i: i["index"])
assert len(items) == len(rows), "queue and submitted rows differ"
for row, item in zip(rows, items):
    note = item["error"]["code"] if item["error"] else "-"
    print(row["sku"], item["status"], item["run_id"], note)

A checklist before you trust the join

Freeze the row list before you build the request. If a later step drops a row with a missing image, drop it before the request is built, not after, because on Sume the item list is the identity of each row. Keep the submitted list next to the queue id in your own storage.

Assert the lengths match when you read the queue. A queue with fewer items than rows means the create call was built from a different list than the one you are joining against. The snippet above does this with a single assert, and that assert is the cheapest data-integrity check in the pipeline.

On the OpenAI side the equivalent habit is to generate custom_id from a stable business key and to treat a missing id in the output as a failure, not as a row to skip. The batch guide also says output files are deleted 30 days after the batch completes, so copy what you need into your own storage; Sume media URLs, by contrast, are durable media.sume.com links.

Failures are rows too

OpenAI records failed requests separately from successes, so an output file alone does not tell you which inputs were lost. On Sume a failed item still occupies its index. When the child could not start, run_id is null and error carries the create-run failure; when the child ran and failed, the reason is on the run receipt at GET /v1/format-runs/{run_id}, not only on the queue item.

Queue status completed means every item is terminal, not that every item succeeded. Branch on counts.failed and counts.canceled before you publish. For the retry side, see bulk run completed is not succeeded.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume