Retry only the failed episodes from a Sume bulk queue

A Sume bulk queue shows completed even when items failed. Find the failed children and re-queue only those, with fresh keys and untouched keys for the rest.

4 min readSume
All posts

Read counts.failed on the queue, list the failed children, and send a new bulk queue containing only those items under new item-level versions. Do not resubmit the whole batch: the finished episodes would either replay for free under their old keys or, if you changed the keys, be paid for twice.

The trap is the queue status. A queue becomes completed when nothing is left to run, not when everything worked. The counts object separates completed from failed and canceled (Sume docs: Bulk runs, read 2026-10-06).

Find what failed

Poll GET /v1/format-run-queues/{id} until the status is completed, then read counts. If failed or canceled is above zero, fetch each child run at GET /v1/format-runs/{run_id}. The item error codes are format_run_failed, format_run_canceled and format_run_failed_to_start. The last one means the child never started, so there is little to investigate, and a retry is usually safe.

For a failed run, read error on the receipt before retrying. A cap that was too low, a bad input or a provider error each call for a different fix. Re-queuing a run that failed because your input was wrong simply buys the same failure again.

Re-queue with the right keys

The queue idempotency key is per Format, and a replay returns 202 with the old queue. So the retry queue needs its own queue key. For each retried item, bump its version in the item key, for example e3-v1 to e3-v2. Items that already succeeded keep their keys and are simply not in the new queue.

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
API = "https://api.sume.com/v1"
queue_id = os.environ["QUEUE_ID"]
items = requests.get(API + "/format-run-queues/" + queue_id,
                     headers=H, timeout=30).json()["data"]
print(items["status"], items["counts"])
if items["counts"]["failed"] == 0:
    raise SystemExit("nothing to retry")
print("Fetch each child, fix the cause, then POST a new")
print("bulk-runs queue with only those items and bumped keys.")

Guard the retry

Set concurrency low for a retry, since failures often come from a shared cause such as the workspace's generation concurrency. Keep generation_spend_cap_usd on every item. There is no cancel-queue endpoint, so if the retry also starts failing, cancel the children one by one with the run cancel endpoint, which returns cancel_effect canceled or no_op.

Retry decision by item error, read 2026-10-06 against Sume docs
Item errorMeaningRetry?
format_run_failed_to_startChild never startedYes, same input
format_run_failedStarted, then failedAfter reading error
format_run_canceledCanceled by youOnly if you intend it

Limits of the retry queue

A bulk queue takes 1 to 100 items and a concurrency of 1 to 16, and each item has the same body as a single Format run. A bad item fails the whole create with 400 invalid_request and an index in details, before any queue exists, so a typo in one retried episode will not leave you with half a queue. Each item also takes its own idempotency_key in the body, which is where the version bump goes (Sume docs: Bulk runs, read 2026-10-06).

Because the controller starts every item with on_active_run: "allow", the only brake inside the queue is the concurrency window plus the workspace generation concurrency. For a retry, a window of 1 or 2 is a reasonable start: failures that came from a shared cause are then visible after one or two items rather than after sixteen.

Why a blanket re-run is the expensive mistake

It is tempting to throw the whole list back into a new queue and let the server sort it out. With stable item keys, the finished items would replay and cost nothing, but that depends on the keys matching exactly and the bodies being identical. One edited field and the replay becomes a conflict or, with a new key, a second paid render.

Selecting only the failed items makes the intent explicit and the bill predictable. Write the list of retried episode numbers into your log, and compare it with counts.failed plus counts.canceled. If they do not match, stop and look before you spend.

Finally, remember that items in a queue run with on_active_run set to allow regardless of what you might expect from single runs, so a retry queue will start all of its items at once up to the concurrency window.

A short checklist

Before the retry, confirm the queue status is completed, list the child ids with a failed or canceled outcome, read each error, fix the cause, bump the item version, and send a new queue with a low concurrency and a cap on every item. After it finishes, compare the final set of delivered episodes with the list you originally wanted.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume