Sume queue_full 429 on a Shorts batch: how do I retry it?

queue_full means your accepted-job capacity is full: 6 on Free, 24 on Pro. Wait for jobs to finish, then resubmit with the same Idempotency-Key.

4 min readSume
All posts

429 queue_full means the workspace already holds its full accepted capacity of jobs that have not started: 5 queued plus 1 running on Free, 24 on Pro. Wait for some to finish, then submit again.

It is a different error from 429 rate_limited, which is about request volume, and from 402 insufficient_credits, which is about balance.

Three errors, three fixes

Sume separates four controls: processing concurrency, queue capacity, submit rate limits and balance. Concurrency is a dispatch limit, so a busy workspace still accepts jobs as queued until the queue is full. Only then does the submit fail.

For a batch, that gives you a rule: never submit more than your accepted capacity at once.

Admission errors (read 2026-10-07)
Status and codeMeaningAction
429 queue_fullAccepted capacity fullWait, then resubmit
429 rate_limitedToo many submit requestsBack off, keep the idempotency key
402 insufficient_creditsBalance too lowUpgrade the plan or submit a cheaper request, then retry the same key

A retry loop

Keep a counter of jobs in flight, poll the job status, and submit the next one only when the counter is under your plan's accepted capacity. Use one idempotency key per episode so a retry does not make a second paid job.

If a batch is large, queue Format runs instead of raw jobs, as in the bulk post, and read the Free-plan post for a wave plan. See the admission docs for the table of limits.

  • Submit no more than accepted capacity.
  • Reuse the key on retry, change it for a new episode.
  • Poll status instead of resubmitting blindly.

Instrumenting the batch

Log three numbers per minute while a batch runs: jobs submitted, jobs terminal, and jobs refused with queue_full. If the refused count stays above zero, your submit rate is higher than your capacity and you can slow the loop. If terminal rises while submitted holds still, you have room to submit more.

Use the job status route for the poll, and webhooks where the endpoint supports them, so you do not hammer read routes. The docs treat read limits as poll backpressure, separate from generation concurrency, so a heavy poll can hit its own limit.

Keep the idempotency key with the episode record. When a submit times out, retry with the same key, and you cannot double-pay for the same episode.

Plan-size cheat sheet

On Free, never have more than 6 jobs outstanding. On Pro, 24. On Startup, 48, and on Scale, 120. Count jobs, and remember that an episode may be several jobs.

If you routinely hit the limit, change the plan or change the pattern. A queue of Format runs lets Sume manage the window, which is simpler than your own loop.

If you drive many episodes from one script, put the retry logic in one function and give it a ceiling. Three attempts with growing waits is enough; beyond that, stop and look at the plan limits, since a loop that never ends only fills the queue again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume