fal retry budgets and 1-hour grace vs Sume job error categories
fal's changelog lists per-condition retry budgets and termination grace of up to one hour. Sume job errors carry a category, a next action and retry-after.

fal's changelog lists per-condition retry budgets and a termination grace of up to one hour, while Sume expresses retry advice on the failed job itself: a category, whether it is retryable, a retry-after value and a next action. A client for either API needs the same discipline, which is to retry only what the error says is safe.
The fal changelog also lists a Platform MCP Server to debug failed requests and a Usage API for cost attribution.
What each side exposes
| Feature | fal changelog | Sume docs |
|---|---|---|
| Retry control | Per-condition retry budgets | Public job error with retryability and retry-after seconds |
| Shutdown | Termination grace up to 1 hour | Cancel only before generation starts; then 409 |
| Debugging | Platform MCP Server for failed requests | Job events and a request id on every error |
| Cost attribution | Usage API | Usage recorded against the member whose key made the job |
Sume job error categories
Failed Sume jobs expose public error metadata: category, stage, retryability, retry-after seconds, public reason and next action. Internal provider payloads are not public fields. The docs list the common categories and the usual response.
| Category | Typical next action |
|---|---|
| validation | Fix input. |
| auth | Check API key and workspace access. |
| quota | Add funds or lower request cost. |
| queue | Retry later with the same idempotency key. |
| generation_unavailable | Retry later. |
| generation_rejected | Inspect events and fix unsupported input. |
| generation_timeout | Poll status or retry later. |
| runtime_unavailable | Retry later; do not retry aggressively. |
| internal | Inspect events and contact support with the request or job id. |
A retry rule that fits both
Retrying blindly turns one failure into several bills. A safer rule is to let the error choose.
- Never retry
validation,authorgeneration_rejectedwithout changing something. - For
queueand capacity errors such asprovider_capacity_exceeded, wait and retry with the same idempotency key. - Use
retry-afterwhen it is present; otherwise back off with a ceiling. - Send an idempotency key on every submit that might be retried; reuse a key only for the same operation and payload.
Do not rely on delivery alone
Webhooks add one more failure to plan for. Sume tries a delivery up to 10 times at a fixed spacing, with a 10-second timeout per attempt. After that you have a failed delivery and a job that still reached its real terminal state.
Keep status polling available for the events that never arrive, and use POST /v1/jobs/{job_id}/webhook/redeliver to re-send a job's real terminal event with a fresh timestamp and signature.
Sources
Related posts
More in Comparisons
- fal's Usage API vs Sume's per-job, per-run cost reads
fal's changelog lists a Usage API for cost attribution. Sume's GET /v1/usage filters by job_id, run_id or thread_id and returns a summary. How to attribute.
- FLUX.2 [klein] needs about 13 GB VRAM: local run or API?
Recraft cites about 13 GB of VRAM for FLUX.2 [klein] under Apache 2.0. A checklist for deciding between a local GPU and a hosted image API.
- Gemini API webhooks for batch jobs: keep a poll backup, as on Sume
Gemini API release notes say webhooks replace polling for Batch and long-running operations. What changes in a client, and why Sume keeps a polling backup.
- Gemini Omni Flash 1,240 vs Seedance 2.0 1,225: what 15 Elo means
Hedra lists Gemini Omni Flash first at 1,240 Elo and Seedance 2.0 4K second at 1,225. A 15-point gap is about a 52% win rate; check limits before choosing.
Written by Sume