Format run failed provider_unavailable or mcp_unavailable: retry rules
provider_unavailable and mcp_unavailable are Sume-side Format run failures: retry with a new Idempotency-Key. provider_credits_exhausted waits.

provider_unavailable and mcp_unavailable are failures on Sume's side of a Format run, not your input, and both are retried with a new Idempotency-Key. provider_credits_exhausted is also Sume's side, but it is not retried right away. All three are in the run-time failure table of the Errors and spend docs.
Treat the set as open: new codes may appear, so handle the ones you know and fall through on the rest.
What does each code mean?
These codes arrive on a run that already exists, as status: "failed" with error: { code, message }, in the run-time failure table of the Errors and spend docs.
| error.code | What happened | Retry? |
|---|---|---|
mcp_unavailable | The per-turn Sume tools the run needed did not attach, so the host failed it before the model ran. details.charged is false | Yes, new Idempotency-Key |
provider_unavailable | The model's provider stream was cut and its reconnects ran out before the deliverable. details.retryable is true | Yes, new Idempotency-Key; finished clips are on the thread and not regenerated |
provider_credits_exhausted | Sume's model provider account ran out of credit. Not your balance. details.retryable is false | Not right away; retry later once Sume reports it restored |
output_extraction_failed | Projection could not run. harvest_unavailable stays completed and fills in on the next read | Read once more, then new key |
Why a new Idempotency-Key?
The old key is bound to the receipt you already have. Replaying it with the same body returns that receipt with idempotency_hit: true and starts nothing, which is exactly what you want after a network blip and exactly what you do not want after a failed run. Derive the key from the thing being made plus an attempt counter, such as order-8823-v1-retry-1.
If the failed run left clips behind, prefer continuing it with previous_run_id over a fresh run, as the docs advise.
How do I encode this in a webhook handler?
Over a webhook the same run arrives with status: "ERROR", outcome: "error" and error.code mirroring the receipt. Dedupe on request_id, which repeats on every retry. Here is a small decision function you can run with Node 18 or later:
const RETRY_NEW_KEY = new Set(["mcp_unavailable", "provider_unavailable"]);
const WAIT = new Set(["provider_credits_exhausted"]);
function decide(run, attempt) {
const code = run.error?.code;
if (!code) return { action: "none" };
if (RETRY_NEW_KEY.has(code)) {
return { action: "retry", key: `order-8823-v1-retry-${attempt + 1}` };
}
if (WAIT.has(code)) return { action: "retry_later" };
return { action: "inspect", code };
}
console.log(decide({ error: { code: "provider_unavailable" } }, 0));
console.log(decide({ error: { code: "provider_credits_exhausted" } }, 0));
console.log(decide({ error: { code: "unattended_blocked" } }, 0));What is billed on a failure?
Generation that finished before a failure or cancel is billed, and a later step failing does not refund it. For mcp_unavailable the docs state details.charged is false: no generation ran. Read usage.billable_amount_usd_micros against usage.generation_spend_cap_usd_micros on the receipt, since a run that wanted to spend past its cap lands on the generic format_run_failed.
A create-time 4xx, an idempotent 200 replay and a skipped run cost nothing. See Sume Format run failure codes for the codes that do need a change on your side, such as unattended_blocked.
How many times should I retry?
The docs give no attempt limit, so choose one and cap your own spend. Retry once with a new key, then contact support with the request_id if the same code repeats, which is the documented advice for the generic format_run_failed_to_start as well. If mcp_unavailable repeats, the docs' reading is that the tools are down, not your request.
Sources
Related posts
More in Formats
- Format run instruction: 8000 characters accepted, about 4000 carried
A Sume Format run instruction accepts 8,000 characters, but only the first ~4,000 reach the agent as prompt text. Put long data in input, carried whole.
- Format run status_url or result_url: which one do I poll?
Poll status_url for a small payload, then read result_url once the run is terminal. result_url answers 409 run_not_completed while the run is in flight.
- Format run stuck in queued: read queue.state and retry_after_seconds
A Sume Format run that stays queued carries a queue block. waiting is normal; runtime_unavailable means nothing claimed it. What to read, and when to ticket.
- OpenAI response_format json_schema to a Sume output_schema
Moving a json_schema from OpenAI structured outputs to a Sume Format run: the field name, what transfers, no JSON mode, and why output comes once, post-run.
Written by Sume