Bulk run: a bad item fails the whole create, a child failure does not

In a Sume bulk Format run, a bad item is a 400 with details.index and no queue; a child that fails admission after the 202 becomes one failed item.

5 min readSume
All posts

A Sume bulk Format run reports problems at two different moments. A bad item, or an attachment that cannot be resolved, fails the whole POST .../bulk-runs create with a 4xx or 5xx and no queue exists. Once the create returns 202, a child that fails to start becomes a single failed item and the rest of the queue keeps going.

Knowing which moment you are in tells you whether to resend the whole batch or read one row. The details are from Bulk runs and Errors and spend.

What fails the create?

Validation runs over every item before the queue is built. The docs say a bad item fails the create with 400 invalid_request and details.index naming the item, and that nothing is dispatched. The same applies when an item's attachments cannot be resolved: invalid_attachment, attachment_not_found, attachment_too_large or attachment_fetch_failed fire before the queue is created, with their usual 400, 413 and 502 statuses.

Create-level limits are concurrency as an integer from 1 to 16 and items with 1 to 100 entries. Each item needs at least one of instruction, input, previous_run_id or attachments, the same rule as a single run.

When a bulk problem surfaces (read 2026-10-03)
ProblemStatusResult
Item names none of instruction, input, previous_run_id, attachments400 invalid_request, details.indexNo queue; nothing dispatched
Attachment unreachable or too large502 or 413No queue; nothing dispatched
Format inactive or API trigger off409No queue
Child fails wallet, concurrency or cap admission202 already returnedThat item is failed; run_id is null
Child run fails while running202 already returnedItem failed; read the run receipt for why

What happens when a child fails after the 202?

Child runs go through ordinary Format-run admission: wallet, workspace generation concurrency and spend caps. A child that cannot start is that item failed, and the queue create has already returned 202. The item keeps its index, shows run_id: null and carries the create-run failure in error, for example format_run_failed_to_start.

The docs say the rest of the queue continues, and the concurrency window refills from the remaining queued items, so one failed child does not stall a batch of 100.

How do I find the failed rows?

Queue status completed only means every item is terminal. Branch on counts.failed and counts.canceled. The command below reads the queue and prints each failed row with its run id; a null id means no run ever started.

QUEUE_ID="frq_..."
curl -sS "https://api.sume.com/v1/format-run-queues/$QUEUE_ID" \
  -H "Authorization: Bearer $SUME_API_KEY" |
  jq '.data.items[] | select(.status == "failed")
      | {index, run_id, code: .error.code}'

Why read the run receipt and not only the queue item?

For a child run that failed while running, the queue item error is generic: format_run_failed with the message The Format run failed. The docs say to read why on the run receipt at GET /v1/format-runs/{run_id}, where the specific code lives, such as unattended_blocked or output_schema_unsatisfied.

Retrying a failed row means a new run with a new Idempotency-Key; the old key is bound to the receipt you already have. Do not resend the whole batch with the same key, because a replay of a spent key returns 202 and the old queue, not a fresh one.

What does the queue not do?

There is no queue-level webhook, no list-queues endpoint and no cancel-queue endpoint. Cancel a child with POST /v1/format-runs/{run_id}/cancel. Mint a fresh key per batch, and keep per-item webhooks if you want push delivery.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume