studio_agent_upstream_unavailable 503: retry, the run keeps going

A Sume 503 studio_agent_upstream_unavailable is a Sume-side outage: retry create with the same Idempotency-Key, and keep polling a run you already hold.

5 min readSume
All posts

A 503 with error.code set to studio_agent_upstream_unavailable is a Sume-side outage, not a problem with your key or your request. Retry the create with the same Idempotency-Key, and if you already hold a run id, keep polling it: the Formats docs say the run is still executing.

The code shows up on four different calls and the right move differs slightly on each, so this page lays out which is which. Everything here comes from the Sume docs for Errors and spend, Calling a Format, Runs and results, Bulk runs and the Action API trigger.

What does the same 503 mean on create versus read?

On create (POST /v1/formats/{handle}/{slug}/runs and POST .../bulk-runs) nothing ran and nothing was charged. The create docs state that a failed create releases its Idempotency-Key, so resending the same key cannot double-bill you and cannot return a stale failure.

On a read (GET /v1/format-runs/{run_id}, /status, /result) the run already exists. The 503 only means the read failed. The docs say the run keeps executing, so abandoning your poll loop does not stop the run or its spend.

studio_agent_upstream_unavailable by call (read 2026-10-03)
CallWhat the docs sayWhat to do
Create a Format runA Sume-side outage. Retry with the same Idempotency-Key.Back off, resend the same key.
Create a bulk queueA Sume-side outage. Retry later.Resend the same key; a replay of a spent key returns the old queue.
Read a runA Sume-side outage, not your key. The run is still executing.Back off and poll again.
Poll a bulk queueRetry later. The queue is still draining.Back off and poll again.
Action API triggerUpstream layer failure; reported as retryable false with next_action contact_support.A bounded retry is reasonable; escalate if it persists.

Why does the envelope say retryable false on the Action trigger?

The Action API trigger page is candid about a mismatch: studio_agent_upstream_unavailable reports retryable: false and next_action: "contact_support" even though it describes an upstream condition. The same page says a bounded retry is still reasonable and to escalate if it persists.

That means a client that obeys retryable literally will give up on the first 503 from the Action route. If you write your own retry policy, special-case this code: allow a small, capped number of retries with growing delays, then stop and raise an alert with the request_id.

The Action page adds one more detail. The route declares 200, 202, 400, 401, 403, 404, 409, 413, 429 and 500 in the OpenAPI document, but the 503 comes from the upstream layer and is not in the declared set. A generated client may not model it, so make sure your generated client's error path handles an undeclared status instead of crashing on parsing.

What should a retry look like?

Keep one idempotency key per intent, not per attempt. The sketch below uses plain fetch and retries only this one status, with a short capped backoff. It resends the same key, so a retry that races a half-succeeded first attempt returns the original run instead of starting a second one.

async function createRun(body: unknown, key: string) {
  const url = "https://api.sume.com/v1/formats/acme/product-promo/runs";
  for (let attempt = 0; attempt < 4; attempt++) {
    const res = await fetch(url, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.SUME_API_KEY}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(body),
    });
    const json = await res.json();
    if (res.status !== 503 || json.error?.code !== "studio_agent_upstream_unavailable") {
      return { status: res.status, json };
    }
    await new Promise((r) => setTimeout(r, 1000 * 2 ** attempt));
  }
  throw new Error("studio_agent_upstream_unavailable after 4 attempts");
}

What if the 503 hits a poll instead?

Do not mark the run failed. A 429 or 503 inside a poll loop is transient, and the docs say that for a bulk queue the same rule applies: back off rather than treating it as a failed queue. The TypeScript SDK's waitForRun already does this for you, absorbing consecutive transient read failures before it gives up.

Keep the run id and the request_id from each failed call. Quote the request_id if the outage outlasts your retry budget, and never paste API keys or signed URLs into a ticket.

What does Sume not promise here?

The docs describe the code as a Sume-side outage and give no duration or status page for it, so there is no number to build a timeout around. They also do not list a retry-after value for this code. Pick your own cap, keep the key stable, and let the run continue while you wait.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume