Free plan accepts 6 paid jobs: submit 7 and read queue_full

Free allows 1 processing job and 5 queued, so 6 are accepted. A Python script submits 7 in parallel and prints which are accepted and which get 429 queue_full.

5 min readSume
All posts

On the Free plan Sume accepts six paid generation jobs at a time: one processing and five queued. A seventh submit, while the other six are still unfinished, returns 429 queue_full. The Python script below submits seven image jobs in parallel and prints the status and error code of each, so you can see the limit before a batch finds it for you.

Every accepted job is real paid work with a reserved balance, so run it only on a workspace and a balance you are happy to spend. Image jobs are used because they are the cheapest generation type to try.

The limits behind the number

The queue capacity default is max(3, concurrency_limit x 5), and accepted capacity is concurrency plus queue. The dashboard Concurrency tab and the generation_limits object in submit responses are the source of truth for your workspace, not this table.

Default processing concurrency and queue capacity by plan (Sume docs read 2026-10-08)
PlanProcessingQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

The script

Each submit runs in a thread through asyncio.to_thread, so the seven requests overlap. The Idempotency-Key per index means a rerun of the script does not bill the same job twice, because a replay returns the original job. Set SUME_BASE only if you point at another host; by default it calls the production API.

import asyncio, json, os, urllib.error, urllib.request

BASE = os.environ.get("SUME_BASE", "https://api.sume.com")

def submit(i):
    req = urllib.request.Request(
        BASE + "/v1/images", method="POST",
        data=json.dumps({"model": "sume/auto", "prompt": f"Plain card {i}", "mode": "async"}).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json", "Idempotency-Key": f"cap-test-{i}"})
    try:
        with urllib.request.urlopen(req, timeout=30) as r:
            return r.status, "accepted"
    except urllib.error.HTTPError as e:
        return e.code, json.load(e).get("error", {}).get("code")

async def main():
    results = await asyncio.gather(*(asyncio.to_thread(submit, i) for i in range(1, 8)))
    for i, (status, code) in enumerate(results, 1):
        print(i, status, code)

asyncio.run(main())

Reading the output

Expect six lines with 202 accepted and one with 429 queue_full. Which index gets the rejection depends on request timing, so do not rely on it being the last one. If the workspace already has other jobs in flight, fewer than six will be accepted.

queue_full is not a rate limit. It means all accepted capacity is used. Sume's guidance is to stop adding work for that workspace, poll current jobs until at least one is terminal, cancel queued jobs you no longer need, and retry with the same idempotency key.

What to do with it

  • Size your worker pool from generation_limits.queue_capacity_remaining in the submit response, not from a constant.
  • Keep the in-flight budget at max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue capacity remaining.
  • Treat queued as healthy. Concurrency limits apply when workers move a job to processing, not when the API accepts it.

Cleaning up after the test

The six accepted jobs are real. Let them finish, or cancel the queued ones with POST /v1/jobs/{id}/cancel while they are still queued. A job that has started generating cannot be canceled, and the cancel call returns 409 for it.

Read GET /v1/jobs?status=queued to list what is waiting before you cancel anything, and stop if the list holds work you did not start.

If your workspace limits differ from the table, the script still shows the truth for your account. Count the accepted lines, and compare the number with the Concurrency tab in the dashboard.

Run the script once, read the six accepted lines, and write the number down. It is the figure your batch runner should not exceed for this workspace until you change plan.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume