Avatar video 429: rate_limited vs queue_full, and which retry to use

Both are HTTP 429 on avatar video submit. rate_limited uses retry-after; queue_full means no accepted capacity. A small Python helper to pick the wait.

4 min readSume
All posts

Two different 429s

The Generation admission docs list two 429 codes. queue_full means the workspace has no remaining accepted generation capacity. rate_limited means request volume exceeded an abuse-protection limit.

429 on a paid submit (Sume Generation admission docs)
CodeMeaningWhat the docs say to do
queue_fullNo accepted capacity leftWait for jobs to finish or cancel queued jobs, then retry with the same idempotency key
rate_limitedToo many requestsUse retry-after for backoff when present, with an idempotency key

Capacity numbers

Plan concurrency and queue capacity from the docs: Free 1 and 5, Pro 4 and 20, Startup 8 and 40, Scale 20 and 100. A queued job counts toward capacity until it starts.

A helper that picks the wait

Run it as is; it needs no network.

def wait_seconds(status, code, headers, attempt):
    if status != 429:
        return None
    if code == 'rate_limited' and 'retry-after' in headers:
        return float(headers['retry-after'])
    return min(60.0, 2.0 ** attempt)

print(wait_seconds(429, 'rate_limited', {'retry-after': '7'}, 1))
print(wait_seconds(429, 'queue_full', {}, 3))

Rules

  • Keep the same Idempotency-Key across retries so a successful earlier attempt is not billed twice.
  • Under queue_full, retrying fast does not help; capacity opens when a running job finishes.
  • Cancel works only before generation starts; after that the API returns 409 job_generation_already_started.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume