Sume 429 concurrency cap or queue cap: which service-account limit?

service_account_concurrency_cap_exceeded counts active jobs; service_account_queue_cap_exceeded counts queued ones. Both are 429. Wait for a job to finish.

4 min readSume
All posts

A 429 from a service-account key can come from two job-count limits, and they count different things. service_account_concurrency_cap_exceeded fires when the account's active generation jobs already equal its concurrency limit. service_account_queue_cap_exceeded fires when its queued generation jobs already equal its queue limit. Both are a count of your own in-flight work, not a request rate, so a delay on the clock is a poor fix. A finished job is the real fix.

The two bodies

Each error puts the live count next to the limit in details, which makes the cause obvious without a second call.

Service-account job-count limits (Sume API source, read 2026-10-05)
Codedetails fieldsCounts
service_account_concurrency_cap_exceededactive_generation_jobs, generation_concurrency_limitActive generation jobs
service_account_queue_cap_exceededqueued_generation_jobs, queued_jobs_limitQueued generation jobs

Why the envelope says retryable

Both return status 429, and the generic 429 branch marks them category: rate_limit, retryable: true. The retry_after_seconds value comes from the response, and the policy code that raises these two passes no retry delay, so retry_after_seconds can be null. Do not invent a delay. Use your own job list instead.

The service-account check is a gate in front of the normal admission rules. A workspace can also queue behind its tier limits; see Generation admission for those, which are a separate layer.

A submit loop that waits on jobs, not on time

Keep a small in-flight set. Submit only when the set is under the limit from the error, and release a slot when a job reaches a terminal state:

import asyncio, json

BODY = '{"error": {"code": "service_account_concurrency_cap_exceeded", "details": {"active_generation_jobs": 3, "generation_concurrency_limit": 3}}}'

async def main():
    err = json.loads(BODY)["error"]
    limit = err["details"].get("generation_concurrency_limit") or err["details"].get("queued_jobs_limit")
    slots = asyncio.Semaphore(limit)
    print("semaphore sized to", limit, "- release a slot when a job is completed, failed or canceled")
    async with slots:
        pass

asyncio.run(main())

Choosing the size of your own window

The error is easiest to avoid when your client already knows the limit. Read generation_concurrency_limit or queued_jobs_limit from the first refusal, store it, and size the in-flight set to it, minus one if other processes share the key. Two workers that each believe they own the whole limit will race, and the loser sees the 429.

Prefer a webhook or a status poll to learn that a slot is free. The jobs and results guide lists the terminal states, and a job that is completed, failed or canceled no longer holds a slot.

Reporting

Show your own users a queue position, not an error. A cap that you know about is a scheduling problem, and the message "3 of 3 slots busy, yours is next" is far better than a failure. Log both numbers from details on every refusal, so you can tell later whether the limit or your own burst size is the thing to change.

When this is the wrong diagnosis

If the code starts with rate_limited or queue_full instead, you are looking at the workspace limiter or admission, which have a retry_after_seconds you can use; the related posts compare them. A cancelled job also stops counting once it leaves the active state, and POST /v1/jobs/:id/cancel exists if you need to drop work you no longer want.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume