Retry by error code, not by 5xx: Sume's 502 is your input

Sume's 502 attachment_fetch_failed means fix your URL, not retry. A retry policy keyed on status and code, with a table of which create errors repeat safely.

6 min readSume
All posts

Do not retry Sume create errors by status class alone. A blanket rule of retry every 5xx and no 4xx gets the most visible case wrong: 502 attachment_fetch_failed is a bad input, and Sume sets its next_action to fix_input even though the status is 5xx. Retry by the pair of status and code. The Formats error page states the principle: branch on the HTTP status first, then on code, and remember that a 4xx at create means nothing ran and nothing was charged.

Which create errors are worth repeating

Only a few create errors repeat safely. 409 idempotency_key_in_use is flagged retryable: true and clears in about a second. 429 rate_limited clears after retry-after. 503 studio_agent_upstream_unavailable is a Sume-side outage, and the docs say to retry later with the same Idempotency-Key. Everything else at create is a statement about your request, and sending it again returns the same answer.

The expensive repeat is 403 insufficient_scope. A key's scopes are fixed when it is minted, so no number of retries will make a formats:read key able to create runs; the docs call looping on it the most common and most expensive mistake. Likewise 402 insufficient_credits returns the same answer until someone tops up in the dashboard.

read 2026-10-03
Create errorCauseRetry?
400 invalid_request, unknown_parameter, output_schema_invalidFix the bodyNever as-is
401 unauthorizedFix the header; two credentials at once also failsNever as-is
402 insufficient_creditsTop up in the dashboardAfter funds only
403 insufficient_scopeMint a new keyNever as-is
409 idempotency_key_in_useAnother request holds the keyWait about 1 s, same key
409 format_run_in_progresson_active_run was rejectWait, or change the option
429 rate_limitedBudget spent, details.scope names itWait retry-after, same key
502 attachment_fetch_failedYour URL is unreachableFix the URL first
503 studio_agent_upstream_unavailableSume-side outageRetry later, same key

The policy in code

The function below encodes that table. The special case for the 502 sits first so the generic rules never see it, and the final fallback follows the docs for format_run_failed_to_start: retry once, then contact support with the request id. It runs as-is.

def retry_policy(status: int, code: str, retryable=None) -> str:
    if status == 502 and code == "attachment_fetch_failed":
        return "fix_input"                      # a 5xx that is your fault
    if status == 409 and code == "idempotency_key_in_use":
        return "wait_1s_same_key"
    if status == 409:
        return "never"                          # conflict, in_progress, inactive...
    if status == 429:
        return "wait_retry_after_same_key"
    if status == 503 and code == "studio_agent_upstream_unavailable":
        return "retry_later_same_key"
    if 400 <= status < 500:
        return "never"                          # nothing ran; fix the call
    if retryable is True:
        return "retry_later_same_key"
    return "retry_once_then_contact_support"

for s, c in [(502, "attachment_fetch_failed"), (409, "idempotency_key_in_use"),
             (409, "format_run_in_progress"), (403, "insufficient_scope"),
             (429, "rate_limited"), (503, "studio_agent_upstream_unavailable")]:
    print(s, c, "->", retry_policy(s, c))

What the SDK already does

If you use the TypeScript SDK, createSumeClient retries 408, 429 and 5xx and transport failures twice with exponential backoff and jitter, honours retry-after, and retries a POST only when an Idempotency-Key is present. That is a bounded safety net, not your policy. A 502 attachment failure will get its two quick retries and then surface as an error; your code still has to read the code and stop. The typed error subclasses help: authentication (401), insufficient credits (402), permission (403), conflict (409) and rate limit (429) each have their own class.

Log the decision

Record the status, code, request id and the policy's verdict on every failed create. A day of those lines tells you quickly whether failures are mostly bad inputs, which means a validation gap in your pipeline, or mostly capacity, which means pacing.

Fixing the attachment cause

A failed fetch usually means the URL is not publicly reachable: a signed link that expired, a private bucket, a host that blocks unknown clients, or a redirect to a login page. details.index names which attachment failed, so fix that one and leave the others. Keep attachment URLs valid for longer than your slowest retry loop, and prefer uploading the file so the Format receives an asset id. Two sibling errors are easy to confuse with it: 400 invalid_attachment means the item itself is malformed or the shared media budget is exceeded, and 413 attachment_too_large means an image over 30 MB or a set over 500 MB. None of the three is fixed by a retry.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume