Avatar batch: which limit hits first, writes per minute or the queue?
On Pro, 300 writes/min is far above the 24 accepted jobs (4 running, 20 queued). Queue capacity limits an avatar batch first, so submit in waves of that size.
For an avatar batch the accepted-job capacity of your plan runs out long before the writes-per-minute rate limit does. On Pro, the rate limit allows 300 writes per minute, but only 24 jobs can be accepted at once (4 running plus 20 queued), so submit in waves of that size and wait for jobs to finish before you add more.
The two limits
Sume has two separate controls. The rate limit counts API requests per minute and returns 429 with a rate-limit signal. The admission limit counts jobs: queued is a normal state, and 429 queue_full fires only when accepted capacity is full. Reads count against a separate, larger bucket: 40 times the write limit.
| Plan | Writes per min | Reads per min | Running | Queue | Accepted |
|---|---|---|---|---|---|
| Free | 120 | 4800 | 1 | 5 | 6 |
| Pro | 300 | 12000 | 4 | 20 | 24 |
| Startup | 600 | 24000 | 8 | 40 | 48 |
| Scale | 1200 | 48000 | 20 | 100 | 120 |
What this means for 200 clips
A Pro workspace can send 200 submit requests in under a minute without hitting the write limit, but about 176 of them will be rejected with queue_full, since only 24 fit. A wave sized to the accepted capacity works better: submit 24, poll with next_poll_after_seconds, and top up as jobs reach a terminal state. Polling is a read, and reads have 40 times the room.
def refill(pending, active, accepted=24):
"""Return the next jobs to submit given counts."""
room = max(0, accepted - len(active))
return pending[:room], pending[room:]
batch, rest = refill(list(range(200)), active=[1, 2, 3])
print(len(batch), len(rest)) # 21 179A wave loop in practice
The loop is simple: keep a list of clips to send, keep a set of job ids that are not terminal, and on every poll cycle send as many new jobs as there are free slots. When a job reaches a terminal state (completed, failed, or canceled), remove it from the set. A failed job is not a reason to resubmit the whole batch; look at its error, fix the script, and queue that clip once with a new key.
Because rendering takes longer than a request, the loop spends nearly all its time waiting. A 200-clip batch on Pro is 9 waves of 24 at most, and the total time is set by how long each wave takes to render, not by how fast you can send requests. If you need it faster, a higher plan has more running slots, which shortens the number of waves.
Practical rules
- Track your own count of non-terminal jobs; do not rely on retrying
queue_full. - Use an
Idempotency-Keyper clip so a retry after a timeout does not make a second paid job. - Prefer webhooks for completion and use polling as a backup; both avoid wasteful reads.
- Check the plan on your account for the current figures before you size a wave.
Limits of this advice
The docs do not say how admission is shared with other job types, so leave headroom. The numbers above are from the docs on the date read, and they may change.
Sources
Related posts
More in Developers
- Avatar create 400: removed name and file fields and their replacements
Old avatar model-run requests that send name or file return 400 with details.fields listing replacements: avatar_handle and input.image_url. Fix both at once.
- 409 avatar_handle_reserved: why handles starting sume_ are off limits
Creating an avatar whose handle starts with sume_ returns 409 avatar_handle_reserved. What the error body says and how to pick a handle that passes.
- Avatar handle rules: 2 to 30 characters, with a Python check
A Sume avatar_handle allows a-z, 0-9, dot and underscore, 2 to 30 characters, no edge or doubled separators. A short Python check, plus the reserved prefix.
- 409 avatar_handle_taken, and how a failed avatar frees its handle
avatar_handle_taken means the handle is in use. When an avatar creation fails at reservation, the old handle is renamed so you can reuse it. Code-verified.
Written by Sume