Unattended agent: stop or retry on Sume 402, 409, 429, 503?
An unattended Sume agent should stop on 402, fix on 400, retry 429 and 503 with the same idempotency key, and never reuse a key after a 409.

An agent with no human watching needs a written rule for each failure, because a model left to improvise will usually retry. The Sume admission page gives clear answers: stop on 402 insufficient_credits, correct the request on 400, wait and retry the same call on 429 and most 503s, and treat 409 idempotency_conflict as a bug in how keys are generated.
The common thread is the idempotency key. A retry that reuses the key for an exact repeat is safe; a retry with a new key after an ambiguous failure can create a second paid job.
The decision table
From the Generation admission page, read 2026-10-09.
| Status and code | Meaning | Agent action |
|---|---|---|
400 invalid_request | Bad body, model id shape, mode, webhook option or header | Fix the request; do not retry unchanged |
401 unauthorized | Key missing, malformed or revoked | Stop and alert a person |
402 insufficient_credits | Balance cannot reserve the estimate | Stop, or submit something cheaper; do not invent top-ups |
404 model_not_found | Model not in this workspace | Check /v1/catalog ids |
409 idempotency_conflict | Same key, different payload | Never reuse a key for a different operation |
429 queue_full | No accepted capacity left | Wait for jobs or cancel queued ones; retry with the same key |
429 rate_limited | Request volume limit | Back off, honor retry-after |
503 provider_capacity_exceeded | Cannot start work safely | Retry later with the same key unless told not to |
Rules to put in the prompt or the wrapper
Write these as code where you can; a prompt is a weaker place for a hard limit.
- Cap total attempts per task, for example three, then report instead of looping.
- Generate the idempotency key once per intended job and keep it with the task, not per attempt.
- On
402, end the run and report the balance; never switch to a more expensive path to compensate. - Treat
queuedas normal. Store the job id and poll with backoff; it is not a failure. - Count a retry against the spend limit you set, since a retried job that did start still counts.
What to log
Log the status, the error code, the request id, the job id and the attempt count. Do not log API keys, signed URLs or raw private media URLs. The Safe automation page lists these as unsafe to log. With that record a person can tell a capacity pause from a billing stop in one look, which is the difference between waiting an hour and topping up a wallet.
Why the boundary is the key
Concurrency being full is not an error. The docs say it becomes a submit error only when the queue is also full. So an agent that reads every slow start as a failure will cancel and resubmit healthy work. Teach it to read generation_limits first: the counts show active jobs, queued jobs and remaining capacity before any retry decision.
Sources
Related posts
More in Agents
- Wan 3.0 clip lengths that fit a $1.00 scheduled run cap
A schedule without its own cap gets $1.00 per run. At Wan 3.0 rates that buys 16 s at 480p, 8 s at 720p or 4 s at 1080p; here is the table.
- Your agent calls generate_video, or Sume's agent does: two meters
With hosted MCP your model client runs the loop and you pay its provider; with an Agent Completion Sume's agent runs it and you set generation_spend_cap_usd.
- Planlock in front of Sume MCP: approve the plan, then the calls
Planlock is an MCP proxy that enforces a human-approved plan. Where it fits in front of Sume's hosted MCP, and which Sume gates still do work behind it.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
Written by Sume