Dagster asset retries and Sume: stop a retry from billing twice

A Dagster RetryPolicy re-runs the whole asset function. Give the Sume submit a stable Idempotency-Key and fail terminal jobs with allow_retries=False.

5 min readSume
All posts

A Dagster retry_policy re-runs the entire asset function, so a paid Sume submit inside it runs again on every retry. To keep that from billing twice, send a stable Idempotency-Key built from your own business key (a SKU and a version, not a run id or a timestamp), and raise dagster.Failure(allow_retries=False) when Sume reports a terminal failure so the policy only retries transport trouble.

Dagster's asset API documents retry_policy as the retry policy for the op that computes the asset, with RetryPolicy(max_retries=3, delay=5) as its usual shape (read 2026-10-04). The Sume side comes from Jobs and results: submit with mode: "async", poll /v1/jobs/{id}/status until terminal is true, then read /result.

Why does the retry policy double-bill without a key?

Dagster retries the op, not the HTTP call. If the first attempt submitted a job and then your process died while polling, attempt two submits again. Without an Idempotency-Key that is a second paid job. With the same key, Sume returns the original job instead of creating another, which is the property you want from a retry.

Two derivations to avoid: a key built from the Dagster run id changes when someone re-materializes the asset, and a key built from a timestamp changes on every attempt. Build it from what the asset means, such as hero-sku-8823-v1, and bump the suffix on purpose when you want a fresh render.

What does the asset look like?

This asset submits one image job, polls with a deadline, and returns the artifact URL.

import os, time, requests
from dagster import asset, Failure, RetryPolicy

API = "https://api.sume.com"
HEAD = {"x-api-key": os.environ["SUME_API_KEY"]}

@asset(retry_policy=RetryPolicy(max_retries=3, delay=5))
def hero_image() -> str:
    r = requests.post(f"{API}/v1/image-1.0/generate", timeout=30,
        headers={**HEAD, "Idempotency-Key": "hero-sku-8823-v1"},
        json={"prompt": "Matte black bottle on marble", "mode": "async"})
    r.raise_for_status()
    job_id = r.json()["data"]["request_id"]
    deadline = time.monotonic() + 20 * 60
    while time.monotonic() < deadline:
        s = requests.get(f"{API}/v1/jobs/{job_id}/status", headers=HEAD, timeout=30)
        s.raise_for_status()
        d = s.json()["data"]
        if d["terminal"]:
            break
        time.sleep(max(d["next_poll_after_seconds"] or 2, 2))
    else:
        raise TimeoutError(f"still running: {job_id}")
    if d["sume_status"] != "completed":
        raise Failure(f"{job_id} ended {d['sume_status']}", allow_retries=False)
    res = requests.get(f"{API}/v1/jobs/{job_id}/result", headers=HEAD, timeout=30)
    return res.json()["data"]["result"]["artifacts"][0]["url"]

Which failures should Dagster retry?

Retry what is transient, stop on what a second attempt cannot fix. The table maps Sume's documented responses to a decision.

Retry decision for a Sume call inside a Dagster asset, from Sume docs read 2026-10-04
ResponseMeaningRetry in the asset?
Timeout or connection reset on submitUnknown whether the job existsYes, same Idempotency-Key
429 rate_limitedRequest budget spentYes, after retry-after seconds
429 queue_fullConcurrency plus queue capacity is fullYes, later, same key
402 insufficient_creditsWallet cannot fund the jobNo, add funds first
Job status failed or canceledTerminal outcomeNo, raise Failure with allow_retries=False
Polling deadline passedThe job keeps running and billingRaise, then resume from the job id

What should you watch out for?

  • A deadline in your poll loop does not cancel the job. It keeps running and billing, so log the job id before you raise.
  • Honor next_poll_after_seconds as a floor; the loop above never polls faster than 2 seconds.
  • Keep the key in the asset's own definition or config, not in a fresh uuid4() per attempt.
  • For minutes-long video, a webhook avoids holding an op open; see Webhooks.

How does this scale to many assets?

A partitioned asset gives you one key per partition for free: use the partition key as the business key, and a backfill of 200 partitions becomes 200 distinct, individually replayable submits. Re-running a failed partition with the same key then returns the job Sume already holds, which is why bumping the version suffix is the explicit way to ask for a new render.

Mind the budget while you fan out. Writes and reads are counted separately per key by plan, so a backfill's submits spend the write budget and its polls spend the much larger read budget. If many partitions run at once, cap Dagster's run concurrency rather than relying on 429 retries, and treat queue_full as a signal to slow the whole backfill, since generation admission limits how many jobs your workspace can hold at once.

What next?

Read the retry rules in Errors and rate limits, then decide per asset whether polling or a signed webhook suits the job length.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume