Bulk ad queue says completed: read counts.failed before you ship

A Sume bulk queue is completed when every item is terminal, not when every ad worked. A Python poller that backs off, then lists failed and canceled items.

5 min readSume
All posts

When GET /v1/format-run-queues/{id} returns status: "completed", every item in the queue has reached a terminal state. It does not mean every ad was made. The Bulk runs page says so twice: completed is not all succeeded, so branch on counts.failed and counts.canceled. A poller that stops at completed and publishes everything will publish the gaps.

What the counts hold

The queue receipt has counts with total, queued, running, completed, failed and canceled, all required. Each row in items has an index, a status, a run_id and an error. A failed item that never started keeps its index with run_id: null and an error such as format_run_failed_to_start. For a child that did run, the item error says only that the run failed. The reason is in the run receipt at GET /v1/format-runs/{run_id}.

Queue item outcomes and what to do (Sume docs, read 2026-10-05)
Item statusrun_iderror codeNext step
completedSetnullRead the run's output
failedSet, or null if it never startedformat_run_failed or format_run_failed_to_startRead the run receipt for the cause, then re-queue
canceledSetformat_run_canceledDecide if you canceled it on purpose
running or queuedSet or nullnullKeep polling

A poller that does the right thing

The code backs off up to a minute, and treats 429 and 503 as temporary, as the docs ask. It returns only when the queue is completed, then reports every item that is not completed. Run it with SUME_API_KEY set.

import os, time, requests

BASE = "https://api.sume.com"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def wait_for_queue(queue_id):
    delay = 5
    while True:
        r = requests.get(f"{BASE}/v1/format-run-queues/{queue_id}", headers=HEAD, timeout=30)
        if r.status_code in (429, 503):
            time.sleep(int(r.headers.get("retry-after", delay)))
            continue
        r.raise_for_status()
        queue = r.json()["data"]
        if queue["status"] == "completed":
            return queue
        time.sleep(delay)
        delay = min(delay * 2, 60)

def main():
    queue = wait_for_queue(os.environ["QUEUE_ID"])
    print(queue["counts"])
    for item in queue["items"]:
        if item["status"] != "completed":
            err = item["error"] or {}
            print(item["index"], item["status"], item["run_id"], err.get("code"))

main()

Re-queue only the failures

Build the next queue from the failed indexes, in the same order, from your own list of items. The queue has no webhook and no public cancel endpoint, and there is no list-queues call, so keep the queue id and your item list yourself. Mint a fresh Idempotency-Key for the second queue. A replayed key returns 202 with the old queue and starts nothing, which is the wrong result when you want a retry.

Before you re-queue, read the run receipts of a few failures. A failure from a wallet or concurrency limit is a reason to wait. A failure from a bad input is a reason to fix the item. Re-queueing the same bad input burns another run.

What to do about cost while you wait

Each child run has its own spend cap, set per item with generation_spend_cap_usd. Use it on every item of an ad batch, because the cap is what limits a runaway run, and it is separate from the queue. If a child cannot start because of wallet or admission limits, that item becomes failed, the create already returned 202, and the window refills from the items still queued.

One more habit helps. Log the queue id, the Idempotency-Key you used and the time of the create next to your own batch id. If your process dies while the queue drains, nothing is lost on the Sume side, because the queue is server-side, and you can resume by polling the id. Without the id, you cannot resume, since there is no list-queues endpoint.

Last, read the queue at a sensible pace. Each child of a video Format is minutes of work, and a poll every second only spends the read budget, which is separate from the create budget. A gap that doubles from five seconds to one minute is enough for a dashboard, and the docs recommend exponential backoff for production.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume