Which Sume submit errors cost money: 402, queue_full, 429, 503
None of them charges you. 402 fails before the reserve, queue_full releases the hold, and rate_limited and 503 mean the job was never accepted.

No submit-time rejection on Sume leaves you with a charge. 402 insufficient_credits fails before provider work starts, 429 queue_full releases the reservation of the failed admission, and 429 rate_limited and 503 mean the job was not accepted. What differs is what you should do next: pay, wait, or back off.
The four errors side by side
The table combines the Generation admission and Errors and rate limits pages.
| Error | Why | Money effect | Do this |
|---|---|---|---|
| 402 insufficient_credits | The estimated cost cannot be reserved from the balance | Fails before provider work | Add funds or submit a cheaper request |
| 429 queue_full | Processing and queue capacity are both full | Reservation of the failed admission is released or refunded | Wait for jobs to finish or cancel queued ones, then retry with the same key |
| 429 rate_limited | Request volume over an abuse limit | Job not accepted | Wait for retry-after, back off, keep the key |
| 503 provider_capacity_exceeded | Dispatch queue is full | Job not accepted | Retry later with the same key |
What the docs say you pay for
Paid generation reserves the estimated amount at submit. A successful completion captures it, and failed jobs or failed queue admission release or refund it. The Format errors page adds that a 4xx at create, an idempotent 200 replay and a skipped run cost nothing, but generation that finished before a later failure is not refunded.
Reading the ledger
If a balance looks lower than expected right after a failed submit, check GET /v1/usage. Reserved rows are holds, not spend, and refunded rows show a hold that was given back. Ask support with the request id from the error body; the docs say it is safe to share, unlike API keys or signed URLs.
- Use the same
Idempotency-Keyfor every retry of one logical job. - Do not retry a paid submit blindly after a local timeout.
retry-afteris the first thing to obey on a 429.
A small decision rule
You can reduce the table to one rule. If the error is 402, change the money. If it is queue_full, change the pace. If it is rate_limited or 503, wait and retry. None of these needs a new idempotency key, and none of them needs a refund request.
For a batch of 100 jobs on a Pro workspace, expect queue_full after 24 accepted jobs if you submit them all at once. Submitting in waves, using the live generation_limits headroom, avoids the error entirely. The admission page recommends max(0, concurrency_limit - active - queued) as the budget for new in-flight work.
- Poll with exponential backoff, not tight loops.
- Cancel queued jobs you no longer need; cancellation succeeds only before generation starts.
- After generation starts, cancel returns
409 job_generation_already_started.
Sources
Related posts
More in Pricing
- A 4-second avatar clip on Sume: $0.736, $0.98 or $2.20 by tier
The shortest accepted Sume avatar clip is 4 seconds: 4 x $0.184 is $0.736, shown as $0.74; Plus is $0.98 and Max is $2.20. Sub-cent rates round up.
- How Sume pricing works: plans, one wallet, published model rates
Sume plans set access and concurrency. Usage draws from one prepaid wallet at each model's published USD rate, for generation, the Agent, Formats, and the API.
- AI avatar video API pricing: cost per second and per minute
Sume bills AI avatar video per second by quality tier, with separate rates when you send a product image. Per-minute costs for standard, plus, and max.
- AI video generation cost per video: what one Sume API run cost
To see what one Sume run or video job cost, call GET /v1/usage with run_id or job_id and read debited_usd: the wallet deduction, agent turns included.
Written by Sume