Free plan accepts 6 paid jobs: submit 7 and read queue_full
Free allows 1 processing job and 5 queued, so 6 are accepted. A Python script submits 7 in parallel and prints which are accepted and which get 429 queue_full.

On the Free plan Sume accepts six paid generation jobs at a time: one processing and five queued. A seventh submit, while the other six are still unfinished, returns 429 queue_full. The Python script below submits seven image jobs in parallel and prints the status and error code of each, so you can see the limit before a batch finds it for you.
Every accepted job is real paid work with a reserved balance, so run it only on a workspace and a balance you are happy to spend. Image jobs are used because they are the cheapest generation type to try.
The limits behind the number
The queue capacity default is max(3, concurrency_limit x 5), and accepted capacity is concurrency plus queue. The dashboard Concurrency tab and the generation_limits object in submit responses are the source of truth for your workspace, not this table.
| Plan | Processing | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
The script
Each submit runs in a thread through asyncio.to_thread, so the seven requests overlap. The Idempotency-Key per index means a rerun of the script does not bill the same job twice, because a replay returns the original job. Set SUME_BASE only if you point at another host; by default it calls the production API.
import asyncio, json, os, urllib.error, urllib.request
BASE = os.environ.get("SUME_BASE", "https://api.sume.com")
def submit(i):
req = urllib.request.Request(
BASE + "/v1/images", method="POST",
data=json.dumps({"model": "sume/auto", "prompt": f"Plain card {i}", "mode": "async"}).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json", "Idempotency-Key": f"cap-test-{i}"})
try:
with urllib.request.urlopen(req, timeout=30) as r:
return r.status, "accepted"
except urllib.error.HTTPError as e:
return e.code, json.load(e).get("error", {}).get("code")
async def main():
results = await asyncio.gather(*(asyncio.to_thread(submit, i) for i in range(1, 8)))
for i, (status, code) in enumerate(results, 1):
print(i, status, code)
asyncio.run(main())Reading the output
Expect six lines with 202 accepted and one with 429 queue_full. Which index gets the rejection depends on request timing, so do not rely on it being the last one. If the workspace already has other jobs in flight, fewer than six will be accepted.
queue_full is not a rate limit. It means all accepted capacity is used. Sume's guidance is to stop adding work for that workspace, poll current jobs until at least one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key.
What to do with it
- Size your worker pool from
generation_limits.queue_capacity_remainingin the submit response, not from a constant. - Keep the in-flight budget at
max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue capacity remaining. - Treat
queuedas healthy. Concurrency limits apply when workers move a job to processing, not when the API accepts it.
Cleaning up after the test
The six accepted jobs are real. Let them finish, or cancel the queued ones with POST /v1/jobs/{id}/cancel while they are still queued. A job that has started generating cannot be canceled, and the cancel call returns 409 for it.
Read GET /v1/jobs?status=queued to list what is waiting before you cancel anything, and stop if the list holds work you did not start.
If your workspace limits differ from the table, the script still shows the truth for your account. Count the accepted lines, and compare the number with the Concurrency tab in the dashboard.
Run the script once, read the six accepted lines, and write the number down. It is the figure your batch runner should not exceed for this workspace until you change plan.
Sources
Related posts
More in Developers
- Free, Pro, Startup, Scale: processing seats, queue slots, full hold
Sume's concurrency by plan, queue capacity max(3, 5 x concurrency), accepted job capacity, and the balance reserved if every slot holds a 10 s clip.
- Gate a Format run on TikTok limits: artifact size, length, pixels
Check a Sume Format run's artifacts[] for size, duration and dimensions against TikTok's non-Spark limits before upload. A Python gate under 30 lines.
- gemini-nano-banana-2.1 or google/nano-banana-2.1: which id Sume takes
Google's pricing page names the model gemini-nano-banana-2.1. Sume's image API takes google/nano-banana-2.1 and runs the old nano-banana-2 as 2.1.
- Gemini Omni video references: 3 clips of 3 s each, $1.31 all in
Sume's Omni route takes up to 3 reference clips of 3 s each. Trimming three source clips at $0.02 each plus a 10 s 720p render at $1.25 is $1.31.
Written by Sume