Bulk queue item with run_id null and format_run_failed_to_start

In a Sume Format bulk queue, an item that never started has run_id null and an error like format_run_failed_to_start. The queue goes on; resubmit that row.

5 min readSume
All posts

When a Sume Format bulk queue item shows status: failed, run_id: null and an error such as format_run_failed_to_start, the item never created a child run, so there is no run receipt to read. The error is the create-run failure from that attempt, the item keeps its index, and the rest of the queue carries on. Resubmit only that row in a later queue.

This is how the bulk runs docs describe it. The create call itself already returned 202, so a start failure on one item does not make the whole request fail.

Two kinds of failed item

A failed item can mean two different things, and you read a different place for each. If run_id is set, the child run started and then failed, and the item error is generic, so open the run. If run_id is null, the child never started and the item error already holds the reason.

Reading a failed bulk-queue item, from docs.sume.com/formats/bulk-runs, read 2026-10-09
Item staterun_iditem errorWhere to look
Child failedarun_...format_run_failed: The Format run failed.GET /v1/format-runs/{run_id}
Child canceledarun_...format_run_canceledNothing to fix unless you did not cancel it
Never startednullCreate-run error, for example format_run_failed_to_startRead error.code and error.message on the item

Why an item might not start

Children go through ordinary Format-run admission: wallet balance, workspace generation concurrency and spend caps. A wallet or admission failure on a child after the 202 does not fail the create. That item becomes failed with the create-run error and the window refills from the remaining queued items.

Counting is the first thing to check. counts.total equals the number of items, and completed for the queue means every item is terminal, not that every item worked. After the queue finishes, branch on counts.failed and counts.canceled.

import json

queue = json.loads("""{"data": {"status": "completed",
  "counts": {"total": 4, "completed": 2, "failed": 2, "canceled": 0},
  "items": [
    {"index": 0, "status": "completed", "run_id": "arun_a", "error": null},
    {"index": 1, "status": "failed", "run_id": null, "error": {"code": "format_run_failed_to_start", "message": "x"}},
    {"index": 2, "status": "failed", "run_id": "arun_c", "error": {"code": "format_run_failed", "message": "x"}},
    {"index": 3, "status": "completed", "run_id": "arun_d", "error": null}]}}""")["data"]

never_started = [i["index"] for i in queue["items"] if i["status"] == "failed" and i["run_id"] is None]
print("resubmit rows:", never_started)

Resubmitting just the gaps

Build a new bulk request from the rows whose index you collected, and use a fresh Idempotency-Key. A replayed key with the same payload returns 202 with the old queue and does nothing new, and a replayed key with a different payload gives 409 idempotency_conflict.

  • Limits for the new queue: concurrency 1 to 16 and items 1 to 100.
  • For items that did start and fail, read the run receipt before retrying, so you do not repeat a failure that will happen again.
  • Per-item communication.webhook_url still works; the queue itself has no webhook.

Reading it as a dashboard

A queue of 100 items at concurrency 16 keeps at most 16 children in flight, so a start failure frees a slot and the next queued item starts at once. That means one bad row cannot stall the batch, but it also means you will not notice it unless you look. A simple rule is to poll the queue until status is completed, then list every item where run_id is null, and alert on that list. Poll with backoff; a 429 or 503 on the poll route is transient and the queue keeps draining.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume