Bulk run completed but SKUs failed: retry only the failures

A Sume bulk-run queue is completed when every item is terminal, not when every item succeeded. Read counts.failed, find the failed SKUs and re-queue only those.

5 min readSume
All posts

Because a bulk-run queue turns completed once every item is terminal, completed does not mean every SKU got a clip. Read counts.failed and counts.canceled, collect the index of each non-completed item, and submit only those items again under a new Idempotency-Key.

Everything below comes from the bulk runs docs.

What does a completed queue actually tell me?

The queue status has three values: queued, running and completed. The docs state that completed means every item is terminal, and that you should inspect counts for failures. Each item ends as completed, failed or canceled, and a terminal item frees its slot for the next one.

That matters for a product catalog because a failed item does not stop the queue. The rest keeps going, and you find out at the end, or sooner if you poll counts.

Queue and item states (from the Sume bulk runs docs, read 2026-10-02)
WhereValueMeaning
Queue statuscompletedEvery item is terminal; not all succeeded
Item statusfailedChild run failed, or never started; run_id may be null
Item statuscanceledChild run was canceled
Item errorformat_run_failedGeneric text; the reason is on the run receipt

How do I find which SKUs failed and why?

Items come back in submission order, each with an index, so keep your own list of SKUs in the order you sent them and map back by index. For the reason, the docs say to read the child run receipt at GET /v1/format-runs/{run_id}, not only the queue item, because the queue item error text is generic.

An item that failed before a run started has run_id: null and an error set from the create attempt, for example format_run_failed_to_start. Treat that as a start failure, and look at wallet, spend cap or concurrency before you retry.

How do I retry only the failed items?

Build a new items list from the failed indexes and send it as a new queue. Use a new Idempotency-Key. The docs are explicit that replaying a spent key returns 202 with the old queue, and that the same key with a different payload is 409 idempotency_conflict.

The sketch below assumes items and skus are the lists you originally sent and that q is the final queue object.

import os, uuid, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
URL = "https://api.sume.com/v1/formats/sume/sume-product-commercial/bulk-runs"

def retry_failed(q, items, skus):
    bad = [i["index"] for i in q["items"] if i["status"] != "completed"]
    if not bad:
        return None
    for i in q["items"]:
        if i["index"] in bad:
            print("retrying", skus[i["index"]], i["status"], i["error"])
    r = requests.post(URL, timeout=30, json={
        "concurrency": 3, "items": [items[n] for n in bad]},
        headers={**H, "Idempotency-Key": str(uuid.uuid4())})
    r.raise_for_status()
    return r.json()["data"]["status_url"]

What about spend and cancellation?

Each item can carry its own generation_spend_cap_usd, per the Format call docs, so a runaway item hits its own ceiling rather than the batch's. Child runs go through ordinary admission, including wallet and workspace concurrency, so a start failure partway through a big batch can be a balance problem rather than a bad input.

To stop an item, cancel the child with POST /v1/format-runs/{run_id}/cancel. There is no public cancel for the whole queue, so plan the batch size accordingly: a 100-item queue is committed once accepted, apart from child-by-child cancels.

What does this not cover?

Whether a failed SKU should be retried at all is your call. If a photo is the cause, such as an unreachable URL or a non-HTTPS link, a retry fails the same way. Fix the input first. Also note the queue has no webhook; use per-item webhooks or poll status_url.

A last habit worth adopting is to persist the queue id and the item index to run id mapping as soon as the create call returns. The create response already lists the first window of items with run ids, and later polls fill in the rest. If your worker restarts halfway through, you can resume by polling the saved queue id instead of submitting again, which would create a second queue and spend twice for the same SKUs. Replaying the same key does return the original queue, so keep the key with the batch record too. With that record you can answer, for any SKU, which queue made its clip, which run id holds the receipt and whether it was a retry, which is exactly what you want when a product page and a video disagree.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume