Failed Format runs: a dead-letter table you can retry from

Keep a dead-letter table for failed Sume Format runs: run id, thread id, failure code, retry flag, and a new idempotency key, so retries stay safe.

4 min readSume
All posts

For failed Sume Format runs, keep a dead-letter table with the run id, thread_id, failure code, whether the failure is retryable, and the business record it belongs to. Retry from that table with a new idempotency key, because the input you sent does not round-trip into the output.

The table turns a pile of red statuses into decisions. Some failures are worth an automatic retry, some need a change to your input, and one class needs nothing but an apology to the customer.

Think of the table as the one place an operator can look at three in the morning. If the answer to what failed and may I retry fits on one row, the integration is in good shape.

What to capture

On a terminal failed receipt you get a failure code and, for provider problems, a retryable indicator. Store both, along with the run id, thread_id, the Format handle and slug, your record id, and the key you sent. The codes are listed on Format errors.

Include the usage block as well. It reports billable_amount_usd_micros for generation, debited_usd_micros for the real wallet deduction including the language model turn, and whether the amount was held, refunded, or final. Finance will want those numbers next to every failure.

Which codes deserve an automatic retry

Provider-side trouble, such as provider_unavailable, is a candidate for a delayed retry. provider_credits_exhausted is marked not retryable, so retrying it only repeats the failure; fix the account side instead. mcp_unavailable is not charged. Failures like output_schema_unsatisfied and deliverable_missing point at the recipe or the schema, so look at the package or the input before retrying.

format_run_failed is the general code and is also used when the generation cap was exceeded, so check usage before deciding it is random.

Suggested handling by failure code, based on the Sume errors page (read 2026-10-10)
CodeRetry?First thing to check
provider_unavailableYes, with delayWait, then new key
provider_credits_exhaustedNoAccount credits
mcp_unavailableYes, not chargedShort delay
output_schema_unsatisfiedAfter a fixSchema versus recipe
deliverable_missingAfter a fixPackage and input
format_run_failedDependsUsage versus cap

Retry with a new key

The Idempotency-Key is scoped to one Format and holds up to 255 characters. The same key and body replays the original run with a 200 and idempotency_hit true. A failed create releases the key, but a run that was created and then failed is a finished run, so to try again you send a new key. Derive it from your record id plus an attempt number, for example order-1042-attempt-2.

Using previous_run_id instead continues the same thread, which creates a new run with its own cap and webhook. That works for revising a result, and it refuses when the earlier run is not terminal or not resumable. Bind the same output_schema again if you use one.

Keeping the table honest

Update a row from the terminal webhook, deduped on request_id, and run a slow reconciler that polls anything still non-terminal after the expiry time. Cap automatic attempts at two or three so a persistent failure cannot become a spending loop. Alert on any row that crosses the limit.

The retry path costs less than people expect when it is narrow. The cookbook's retry-one-scene recipe re-runs only a failed part, and the docs report about one twentieth of the first run's cost in production measurements. Check the Formats cookbook before you re-run a whole video.

Finally, give each row an owner field. A failure that nobody owns is a failure that is retried forever or never. When a person closes a row, store the reason in plain words so the next incident starts with context.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume