Re-queue only failed items of a Sume Format bulk queue

A completed bulk queue is not all succeeded. Read counts.failed, pick the failed indexes, and resend them under a new Idempotency-Key. Python, offline.

5 min readSume
All posts

When a Sume Format bulk queue finishes with failed items, build a new queue from only those items and send it under a new Idempotency-Key. Do not reuse the old key: a replay of a spent key returns 202 with the old queue and starts nothing, and a changed payload under the same key returns 409 idempotency_conflict. The Python helper below reads the queue receipt, picks the failed indexes, and builds the retry body.

A queue with status completed means that every item is terminal, not that every item succeeded. Branch on counts.failed and counts.canceled before you tell anyone the batch is done.

What the receipt gives you

A failed item that never started a run keeps its index with run_id: null and an error such as format_run_failed_to_start. To learn why a started child failed, read its run receipt at GET /v1/format-runs/{run_id}, not only the queue item.

Fields of the format.run_queue receipt used here (Sume docs read 2026-10-08)
FieldNotes
data.statusqueued, running, or completed. Completed means all items are terminal.
data.countstotal, queued, running, completed, failed, canceled
data.items[].indexZero-based position in the items you submitted
data.items[].statusqueued, running, completed, failed, or canceled
data.items[].run_idNull if the item failed before a child run started
data.items[].errorcode and message, or null

The helper

The example queue has five items. Index 1 failed before it started, index 3 failed after it started, and index 4 was canceled. The helper selects only failed, so it returns indexes 1 and 3. Decide separately whether canceled items should be retried, because a cancel is usually deliberate.

import json, uuid

def retry_batch(original_items, queue, concurrency=3):
    failed = [row["index"] for row in queue["data"]["items"] if row["status"] == "failed"]
    body = {"concurrency": concurrency, "items": [original_items[i] for i in failed]}
    return failed, str(uuid.uuid4()), body

items = [{"instruction": f"clip {n}"} for n in range(1, 6)]
queue = {"data": {"status": "completed", "counts": {"total": 5, "failed": 2}, "items": [
    {"index": 0, "status": "completed", "run_id": "arun_a", "error": None},
    {"index": 1, "status": "failed", "run_id": None, "error": {"code": "format_run_failed_to_start"}},
    {"index": 2, "status": "completed", "run_id": "arun_c", "error": None},
    {"index": 3, "status": "failed", "run_id": "arun_d", "error": {"code": "format_run_failed"}},
    {"index": 4, "status": "canceled", "run_id": "arun_e", "error": {"code": "format_run_canceled"}}]}}
failed, key, body = retry_batch(items, queue)
print(failed, len(key), json.dumps(body))

Sending the new queue

Post the body to POST /v1/formats/{handle}/{slug}/bulk-runs with the key from the helper in the Idempotency-Key header. The create accepts a concurrency between 1 and 16 and between 1 and 100 items. Every item must name at least one of instruction, input, previous_run_id, or attachments; a bad item fails the whole create with 400 invalid_request and details.index, before any queue exists.

Keep a map from the new queue's item positions back to your original indexes, because the new queue starts counting from zero. In the example, new item 0 is old item 1 and new item 1 is old item 3.

Cost control

Each item can carry its own generation_spend_cap_usd, and the queue has no queue-level cap. Worst-case spend for a retry is the sum of the caps of the items you resend. Check that sum before you post, and check the balance, since a child that cannot be admitted becomes a failed item after the create has already returned 202.

Deciding what counts as failed

Not every failure deserves a retry. An item that failed with a validation problem in your input will fail again with the same input, so fix the input first. An item that failed to start because of a balance or capacity problem is a good candidate once the cause is gone.

Log each retry as its own queue and keep the link to the first queue id. When a second queue also has failures, you can then see how many rounds an item has been through, and you can stop after two. Retrying forever turns a data problem into a spend problem.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume