Format run provider_credits_exhausted: not your balance, wait
provider_credits_exhausted means Sume's model provider account ran out of credit, not your balance. Do not retry right away; wait, then use a new key.

provider_credits_exhausted is a Format run failure that means Sume's own model provider account ran out of credit, so the run stopped. It is not your Sume balance and not your input, and details.retryable is false. Do not re-fire right away; retry later with a new Idempotency-Key once Sume reports the provider account restored.
This is from the failure table in Errors and spend, which is the only place the docs define it.
How is it different from insufficient_credits?
The two sound alike and fail at different stages. insufficient_credits is a 402 at create: your workspace wallet cannot fund the run, nothing ran and nothing was charged, and next_action is add_funds. provider_credits_exhausted arrives after a 202, on the receipt as status: "failed", and no top-up on your side changes it.
| insufficient_credits | provider_credits_exhausted | |
|---|---|---|
| When | At create, as an HTTP 402 | After 202, on the run receipt |
| Whose balance | Your workspace wallet | Sume's model provider account |
| Retryable | Not until you top up | details.retryable is false |
| What to do | Add funds, then resend | Wait, then retry with a new Idempotency-Key |
| Charged | Nothing ran | Generation finished before the stop is billed |
What should my code do when it sees it?
Stop sending new runs of that kind, and do not enter an automatic retry loop. A tight retry gives the same failure and, because earlier generation is billed, it can add cost. Alert a person, hold the queue and retry after a delay you choose.
Treat the code set as open, as the docs say, and keep the generic path for codes you do not know. Compare it with the Sume-side outage codes: provider_unavailable and mcp_unavailable are retryable with a new key, while this one is explicitly not retryable right now.
How does it look inside a bulk run?
A child run that fails this way is a failed queue item, with a generic format_run_failed code on the queue. The reason is only on the child receipt, so read GET /v1/format-runs/{run_id} for each failed index and group by error.code before you requeue. If many items share provider_credits_exhausted, requeueing them at once will only repeat the failure.
import collections, os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
q = requests.get(os.environ["QUEUE_URL"], headers=H, timeout=30).json()["data"]
codes = collections.Counter()
for it in q["items"]:
if it["status"] == "failed" and it["run_id"]:
r = requests.get(f"https://api.sume.com/v1/format-runs/{it['run_id']}",
headers=H, timeout=30).json()["data"]
codes[(r.get("error") or {}).get("code", "unknown")] += 1
print(dict(codes))
if codes.get("provider_credits_exhausted"):
print("hold: do not requeue yet")Will I be charged for the failed run?
Generation that finished before a failure is billed, and a later step failing does not refund it, so check usage.billable_amount_usd_micros on the receipt. A 4xx at create costs nothing. The docs do not say how long a provider outage lasts or how Sume reports that the account is restored, so build your retry delay as a setting, not a constant.
For the retryable failures see provider_unavailable and mcp_unavailable retry rules.
Sources
Related posts
More in Formats
- Format run status_url or result_url: which one do I poll?
Poll status_url for a small payload, then read result_url once the run is terminal. result_url answers 409 run_not_completed while the run is in flight.
- Format run stuck in queued: read queue.state and retry_after_seconds
A Sume Format run that stays queued carries a queue block. waiting is normal; runtime_unavailable means nothing claimed it. What to read, and when to ticket.
- Format run failed with unattended_blocked: why it never half-finishes
Over the API, a Sume Format run is told approvals are granted. If it still cannot finish it fails with unattended_blocked, never a half-done completed.
- Gemini batch structured output per request vs Sume per-item schema
Gemini batch requests can each carry a JSON schema. A Sume bulk item is a full run body, so each item can bind its own output_schema; read output on each child.
Written by Sume