Sume bulk queue 404 format_run_queue_not_found: three causes

A Sume bulk queue poll returned 404 format_run_queue_not_found. The id is wrong or the queue is another owner's. How to tell which, and what to do next.

4 min readSume
All posts

Why does GET /v1/format-run-queues/{id} return 404 format_run_queue_not_found? The Bulk runs docs give two causes in one line: an unknown queue id, or a queue that belongs to another owner. The API deliberately makes those two look the same, because telling a caller that someone else's queue exists would leak it. A queue you cannot see reads as one that does not exist.

That makes a third cause worth checking first in practice: the key you are polling with is not the key that created the queue.

Check in this order

  • Compare the id with the id in the 202 create receipt. A truncated or edited frq_ id is the most common cause.
  • Confirm the same user owns the key you poll with. A queue made with one person's key is not readable by a teammate's key.
  • Check the environment. api.sume.com and api.dev.sume.com are separate hosts, and a queue made on one is not on the other.
  • Check the scope. A key without formats:read gets 403 insufficient_scope, not a 404, so a 404 is not a scope problem.

Log what you need next time

Write the queue id, the key's owner label, the host and the Format address into one log line at create time. When a 404 shows up days later, those four facts answer which of the three causes you hit within seconds.

If several services share one pipeline, give the poller the same credentials as the creator, or hand the creator's receipt to the poller as data. A separate key for the poller works only if it is the same owner.

Rate limits are separate for reads and writes, so a poller that retries a bad id in a tight loop spends the read budget for nothing. Stop on the first 404.

What the 404 does not mean

It does not mean the queue finished and was cleaned up, and it does not mean the work stopped. The docs say a 429 or 503 during polling is transient and the queue keeps working, but they do not describe a retention window for queues, so do not build logic that treats a later 404 as proof of completion. If you need the outcome long term, store each child run's receipt as the queue drains.

Lost the id? There is no list-queues endpoint. The child runs are ordinary Format runs, so if you saved any run_id, read its receipt at GET /v1/format-runs/{run_id}; the queue itself cannot be rediscovered from it by the documented API.

A poll loop that classifies the 404

The function below separates a hard 404 from transient errors, so a typo stops the loop at once instead of being retried. The HTTP call is a stub.

def classify(status, code=None):
    if status == 200:
        return "ok"
    if status == 404 and code == "format_run_queue_not_found":
        return "stop: wrong id, wrong owner or wrong host"
    if status in (429, 503):
        return "retry with backoff"
    return "stop: read the error body"


for s, c in [(200, None), (404, "format_run_queue_not_found"), (429, "rate_limited"), (503, None), (403, "insufficient_scope")]:
    print(s, c, "->", classify(s, c))

Sources

Related posts

More in Developers

All Developers posts

Written by Sume