fal retry budgets and 1-hour grace vs Sume job error categories

fal's changelog lists per-condition retry budgets and termination grace of up to one hour. Sume job errors carry a category, a next action and retry-after.

4 min readSume
All posts

fal's changelog lists per-condition retry budgets and a termination grace of up to one hour, while Sume expresses retry advice on the failed job itself: a category, whether it is retryable, a retry-after value and a next action. A client for either API needs the same discipline, which is to retry only what the error says is safe.

The fal changelog also lists a Platform MCP Server to debug failed requests and a Usage API for cost attribution.

What each side exposes

Failure handling features named in each source (read 2026-10-03)
Featurefal changelogSume docs
Retry controlPer-condition retry budgetsPublic job error with retryability and retry-after seconds
ShutdownTermination grace up to 1 hourCancel only before generation starts; then 409
DebuggingPlatform MCP Server for failed requestsJob events and a request id on every error
Cost attributionUsage APIUsage recorded against the member whose key made the job

Sume job error categories

Failed Sume jobs expose public error metadata: category, stage, retryability, retry-after seconds, public reason and next action. Internal provider payloads are not public fields. The docs list the common categories and the usual response.

Common job error categories (from the Sume errors page, read 2026-10-03)
CategoryTypical next action
validationFix input.
authCheck API key and workspace access.
quotaAdd funds or lower request cost.
queueRetry later with the same idempotency key.
generation_unavailableRetry later.
generation_rejectedInspect events and fix unsupported input.
generation_timeoutPoll status or retry later.
runtime_unavailableRetry later; do not retry aggressively.
internalInspect events and contact support with the request or job id.

A retry rule that fits both

Retrying blindly turns one failure into several bills. A safer rule is to let the error choose.

  • Never retry validation, auth or generation_rejected without changing something.
  • For queue and capacity errors such as provider_capacity_exceeded, wait and retry with the same idempotency key.
  • Use retry-after when it is present; otherwise back off with a ceiling.
  • Send an idempotency key on every submit that might be retried; reuse a key only for the same operation and payload.

Do not rely on delivery alone

Webhooks add one more failure to plan for. Sume tries a delivery up to 10 times at a fixed spacing, with a 10-second timeout per attempt. After that you have a failed delivery and a job that still reached its real terminal state.

Keep status polling available for the events that never arrive, and use POST /v1/jobs/{job_id}/webhook/redeliver to re-send a job's real terminal event with a fresh timestamp and signature.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume