OpenAI batch output order differs from input: how Sume bulk keys rows
OpenAI batch output may not follow input order, so you join on custom_id. A Sume bulk queue lists items by index in submitted order. Here is the join for each.

Do OpenAI batch results come back in the order you sent them? No. The OpenAI Batch guide (read 2026-10-04) says the lines in the output file may not match the order of the input file, so you map each result to its request with custom_id. A Sume bulk queue works the other way round: the queue receipt lists one item per submitted row, in the same order, each with a zero-based index. Your join key on Sume is the position you sent, not a label you invented.
That difference matters the week a retailer loads a spreadsheet of holiday SKUs and needs each finished video back on the right row. Get the join wrong and a product page shows another product's clip.
The fix is cheap in both systems, but it is a different fix. With OpenAI you carry an identifier through the request. With Sume you carry nothing and trust the order of the receipt, so the one thing you must never do is sort, filter or de-duplicate your rows between building the request and reading the queue.
What each API gives you to join on
The table compares the two join keys as documented. The Sume row comes from the Bulk runs page.
| OpenAI Batch API | Sume bulk run | |
|---|---|---|
| Request label you choose | custom_id on every input line | None. Each item is the body of one Format run |
| Order of results | May not match input order | Same order as submitted items |
| How to join | Match output line to input by custom_id | items[i] belongs to the i-th row you sent |
| Failures | Recorded separately from successes | Item status failed, with error and run_id (null if the run never started) |
| Max size per submission | 50,000 requests, 200 MB input file | 1 to 100 items |
Join a Sume queue back to your rows
Because the queue is ordered, the safest pattern is to keep your rows in a list and zip them with items from the queue. Each item carries index, status, run_id and error. Do not rely on the order in which runs finish; the window starts the next queued item as soon as a slot opens, so completion order is whatever the work dictates. Position in items is stable.
If you want a label anyway, put it in the item's input object, since you choose that shape. The Format reads the keys it knows and ignores the rest, so a sku key costs nothing and makes your own logs readable. That is a convention you add, not a field Sume echoes back as a join key.
A runnable zip
This snippet takes the rows you submitted and a queue receipt you fetched from GET /v1/format-run-queues/{id}, and prints the status per row. It uses only the standard library.
import json
rows = [{"sku": "mug-01"}, {"sku": "mug-02"}, {"sku": "mug-03"}]
queue = json.loads('''{"items": [
{"index": 0, "status": "completed", "run_id": "r0", "error": null},
{"index": 1, "status": "failed", "run_id": null,
"error": {"code": "format_run_failed_to_start", "message": "x"}},
{"index": 2, "status": "running", "run_id": "r2", "error": null}]}''')
items = sorted(queue["items"], key=lambda i: i["index"])
assert len(items) == len(rows), "queue and submitted rows differ"
for row, item in zip(rows, items):
note = item["error"]["code"] if item["error"] else "-"
print(row["sku"], item["status"], item["run_id"], note)A checklist before you trust the join
Freeze the row list before you build the request. If a later step drops a row with a missing image, drop it before the request is built, not after, because on Sume the item list is the identity of each row. Keep the submitted list next to the queue id in your own storage.
Assert the lengths match when you read the queue. A queue with fewer items than rows means the create call was built from a different list than the one you are joining against. The snippet above does this with a single assert, and that assert is the cheapest data-integrity check in the pipeline.
On the OpenAI side the equivalent habit is to generate custom_id from a stable business key and to treat a missing id in the output as a failure, not as a row to skip. The batch guide also says output files are deleted 30 days after the batch completes, so copy what you need into your own storage; Sume media URLs, by contrast, are durable media.sume.com links.
Failures are rows too
OpenAI records failed requests separately from successes, so an output file alone does not tell you which inputs were lost. On Sume a failed item still occupies its index. When the child could not start, run_id is null and error carries the create-run failure; when the child ran and failed, the reason is on the run receipt at GET /v1/format-runs/{run_id}, not only on the queue item.
Queue status completed means every item is terminal, not that every item succeeded. Branch on counts.failed and counts.canceled before you publish. For the retry side, see bulk run completed is not succeeded.
Sources
Related posts
More in Comparisons
- Pika's auto model pick vs Sume Auto and pinned ids
Pika's Sept 17, 2026 relaunch lists 25-plus apps, auto or manual model pick and cheap Seedance. How that compares with Sume Auto and a pinned seedance-2.5.
- Pika API Club: 100+ models at reduced pricing, questions to ask first
Pika's API Club (Aug 5) offers 100+ models at reduced pricing, and the Sep 17 relaunch adds audio models. Seven questions to ask an aggregator.
- LiveTranslate 2.3 s latency vs STT plus TTS
Qwen3.8-LiveTranslate cut average lagging from 2.8 to 2.3 seconds. Sume has no live interpreter, but stt_create and tts_create can build an offline dub.
- Capacity fallbacks: Replicate's model swap vs pinning on Sume
Replicate lists a model falling back to another at capacity. On Sume you pin a model or send sume/auto, and a capacity error is retried with the same key.
Written by Sume