Unattended agent: stop or retry on Sume 402, 409, 429, 503?

An unattended Sume agent should stop on 402, fix on 400, retry 429 and 503 with the same idempotency key, and never reuse a key after a 409.

5 min readSume
All posts

An agent with no human watching needs a written rule for each failure, because a model left to improvise will usually retry. The Sume admission page gives clear answers: stop on 402 insufficient_credits, correct the request on 400, wait and retry the same call on 429 and most 503s, and treat 409 idempotency_conflict as a bug in how keys are generated.

The common thread is the idempotency key. A retry that reuses the key for an exact repeat is safe; a retry with a new key after an ambiguous failure can create a second paid job.

The decision table

From the Generation admission page, read 2026-10-09.

Submit errors and the safe response for an unattended agent
Status and codeMeaningAgent action
400 invalid_requestBad body, model id shape, mode, webhook option or headerFix the request; do not retry unchanged
401 unauthorizedKey missing, malformed or revokedStop and alert a person
402 insufficient_creditsBalance cannot reserve the estimateStop, or submit something cheaper; do not invent top-ups
404 model_not_foundModel not in this workspaceCheck /v1/catalog ids
409 idempotency_conflictSame key, different payloadNever reuse a key for a different operation
429 queue_fullNo accepted capacity leftWait for jobs or cancel queued ones; retry with the same key
429 rate_limitedRequest volume limitBack off, honor retry-after
503 provider_capacity_exceededCannot start work safelyRetry later with the same key unless told not to

Rules to put in the prompt or the wrapper

Write these as code where you can; a prompt is a weaker place for a hard limit.

  • Cap total attempts per task, for example three, then report instead of looping.
  • Generate the idempotency key once per intended job and keep it with the task, not per attempt.
  • On 402, end the run and report the balance; never switch to a more expensive path to compensate.
  • Treat queued as normal. Store the job id and poll with backoff; it is not a failure.
  • Count a retry against the spend limit you set, since a retried job that did start still counts.

What to log

Log the status, the error code, the request id, the job id and the attempt count. Do not log API keys, signed URLs or raw private media URLs. The Safe automation page lists these as unsafe to log. With that record a person can tell a capacity pause from a billing stop in one look, which is the difference between waiting an hour and topping up a wallet.

Why the boundary is the key

Concurrency being full is not an error. The docs say it becomes a submit error only when the queue is also full. So an agent that reads every slow start as a failure will cancel and resubmit healthy work. Teach it to read generation_limits first: the counts show active jobs, queued jobs and remaining capacity before any retry decision.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume