Free plan: 120 writes a minute but 6 accepted jobs, which hits first

Video batches on Sume hit queue_full long before the write rate limit. Per-plan arithmetic for 429 rate_limited versus queue_full, with a calculator.

4 min readSume
All posts

For video batches on Sume, the accepted-job limit trips long before the write rate limit does. A Free workspace can send 120 writes a minute but can hold only 6 paid generation jobs, so the seventh simultaneous submit gets 429 queue_full, even though you are using 5 percent of your request budget. A Scale workspace has 1,200 writes a minute and 120 accepted jobs, with the same shape.

Two limits that share a status code

Both limits answer 429, which is why they get confused. rate_limited means request volume exceeded a window, and the response carries ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after. MDN defines retry-after as either a delay in seconds or an HTTP date, and Sume sends seconds.

queue_full means the workspace used all of its accepted generation capacity. It has nothing to do with requests per minute. It clears when a running or queued job finishes or is canceled, which for a 30-second video is minutes, not seconds.

The ratio per plan

Divide the write budget by the accepted capacity and you see how far apart the two limits sit. Reads have their own bucket at forty times the write number, so polling cannot starve your submits.

Write budget versus accepted jobs per plan (Sume docs, read 2026-10-05)
PlanWrites per minuteReads per minuteAccepted jobsWrites per accepted job
Free1204,800620
Pro30012,0002412.5
Startup60024,0004812.5
Scale1,20048,00012010

A calculator for your batch

Give it a plan and a count and it tells you how many waves you need, how much headroom a minute of writes gives you, and which limit is the binding one. It is plain Python, so you can run it before you spend anything.

A short worked case helps. Say you queue 40 Seedance 2.5 clips of 30 seconds on a Free plan. You can fire 40 POSTs in one minute without touching the 120-write budget, but only 6 are accepted at a time, so 34 of them return queue_full. If your code treats every 429 as a rate limit and sleeps for a fixed 60 seconds, you retry about once a minute and the slots sit idle between finishes. A pacer that submits when a job ends keeps the slots full.

PLANS = {  # writes/min, accepted jobs (concurrency + queue)
    "free": (120, 6), "pro": (300, 24), "startup": (600, 48), "scale": (1200, 120),
}

def plan_batch(plan: str, jobs: int, minutes_per_job: float = 3.0):
    writes, accepted = PLANS[plan]
    waves = -(-jobs // accepted)  # ceiling division
    binding = "queue_full" if accepted < writes else "rate_limited"
    return {
        "plan": plan,
        "waves": waves,
        "binding_limit": binding,
        "submits_in_first_minute_allowed": min(jobs, writes),
        "jobs_accepted_at_once": min(jobs, accepted),
        "rough_wall_minutes": waves * minutes_per_job,
    }

for p in PLANS:
    print(plan_batch(p, 100))

Reading the output

For 100 clips the Free plan needs 17 waves and Pro needs 5. The binding limit is queue_full in every row, because accepted jobs are always far below writes per minute. The wall-clock figure is an input you choose, not a measurement, so replace minutes_per_job with what your own jobs take.

The practical advice is to size the job pacer around accepted capacity and let the rate limiter be a safety net. A pacer that sends one submit per slot that frees will never see rate_limited, and rarely see queue_full.

There is one more place the two limits meet: polling. Reads are a separate bucket at forty times the write number, so a status loop over 24 jobs every two seconds is about 720 reads a minute, well inside Pro's 12,000. You can poll generously and still keep your submit budget untouched, which is the reason the split exists.

What to do on each code

  • 429 rate_limited: wait for retry-after, then resend with the same Idempotency-Key.
  • 429 queue_full: wait for a job to finish, cancel queued jobs you do not need, then resend with the same key.
  • Do not retry an unsafe submit without a key. A retry without one can bill twice.
  • Check error.details.scope on a 429. It names the budget, read or write, that the request spent from.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume