Sume job error category quota or queue vs 402 and 429
A Sume job error category quota means add funds or lower cost; queue means retry with the same key. They sit on the job, apart from 402 and 429 at submit.

Sume reports money and capacity problems in two places. At submit time you get an HTTP error: 402 insufficient_credits or 429 queue_full or 429 rate_limited. After a job exists, a failure can carry a job error category such as quota or queue. The documented actions are: quota means add funds or lower request cost, and queue means retry later with the same idempotency key.
Knowing which place you are looking at tells you whether a job exists and whether anything can be billed.
Where does each error show up?
The Errors and rate limits page lists HTTP status codes and a separate table of job error categories. The generation admission page then explains the submit-time cases in detail.
| Where | Signal | Documented next action |
|---|---|---|
| Submit response | 402 insufficient_credits | Balance cannot cover the estimate; upgrade the plan, wait for included Gen$, or submit a cheaper request |
| Submit response | 429 queue_full | Wait for jobs to finish or cancel queued jobs, then retry with the same idempotency key |
| Submit response | 429 rate_limited | Back off using retry-after when present |
| Failed job | category quota | Add funds or lower request cost |
| Failed job | category queue | Retry later with the same idempotency key |
Why would a job fail with quota after it was accepted?
The docs say a submit creates a durable job when the request is valid, balance can be reserved and the workspace still has accepted-job capacity. So a 402 at submit means no job was created. The docs do not explain every path by which a job later fails with category quota, and this post will not guess at one.
What they do give is the response: treat it as a money problem on the job record, add funds or lower the cost of the request, and then submit again deliberately. Read the public error on the job (GET /v1/jobs/{id}) for the reason and next action rather than relying on the category alone.
What does retrying with the same key do?
The queue category and the queue_full submit error share a recipe: retry with the same Idempotency-Key once capacity opens. The key makes the retry return the original job if one exists, and creates a new one if none does, so you cannot double-bill by retrying.
A different key, or a changed payload under the same key, is not a retry. Reusing a key for a different operation or payload returns 409 idempotency_conflict.
curl https://api.sume.com/v1/jobs/job_123 \
-H "Authorization: Bearer $SUME_API_KEY"
# read the job's public error: category, stage, retryability,
# retry-after seconds, public reason and next actionHow do I tell which case I am in from logs?
Look at the shape first. A submit-time error arrives as the error envelope: error.code, error.message, error.request_id and sometimes error.details. A job error arrives on a job record that has a job id, a status of failed, and public error metadata. If there is no job id, there is no job.
Log both the request id and the job id when you have them. The request id is safe to share with Sume support, and the job id lets you pull events from GET /v1/jobs/{id}/events for the public timeline. Keep API keys, signed URLs and raw media URLs out of tickets and logs you share.
What about the reservation?
Paid generation reserves the estimated amount at submit. Successful completion captures it. Failed jobs and failed queue admission release or refund the reservation where applicable. For a failed admission the error details can include a generation_limits snapshot and job metadata for the attempt.
If you hit 402 repeatedly, check GET /v1/balance and your plan before retrying. The docs say not to invent prepaid top-ups as a fix for concurrency, since top-ups do not raise the processing concurrency limit. See what AI credit overages are for how overages are billed or refused.
Sources
Related posts
More in Developers
- Sume job failed with worker_timeout: poll again or retry?
A Sume job error in the worker_timeout or generation_timeout category means poll status or retry later. runtime_unavailable means retry later, gently.
- Sume job metadata on Kling motion control and H3 Max lip sync
Both Kling 3.0 Motion Control and MiniMax H3 Max Lip Sync accept a metadata object stored with the Sume job request. It is not sent to the provider.
- jobs_result with job_ids: read ok per entry, re-read failed_job_ids
A Sume MCP jobs_result batch returns ok plus value or error per id. One job_not_completed does not fail the rest; re-read partial_failure.failed_job_ids.
- jobs_wait include_results: skip the second read, handle omitted ids
Set include_results true on Sume MCP jobs_wait: completed ids return jobs_result in results[]; ids that do not fit are named in results_omitted.job_ids.
Written by Sume