DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.

A DBOS workflow can submit a Sume job, poll it and survive a process crash, because DBOS records each completed step in a database and resumes from the last one. Put the submit in a @DBOS.step with retries, make the Sume Idempotency-Key come from the order id, and start the workflow under a fixed workflow id so the same order can never run twice.
I ran the sample below with DBOS 3.2.0 and a SQLite system database against a local stand-in for the API. One injected 429 was retried, and a second run of the same program with the same workflow id printed the stored result without sending another create request.
What does DBOS give you that a retry loop does not?
DBOS's step tutorial documents DBOS.step with retries_allowed, interval_seconds, max_attempts, backoff_rate and a should_retry callback, and says that after an interruption a workflow resumes from the last completed step without re-executing finished ones. That second property is a cheap form of exactly-once for your own bookkeeping, but it cannot reach across to Sume. A step can still die after Sume accepted the create and before DBOS stored the answer, which is why the key still matters: replaying the same key and body returns the original job with idempotency_hit: true.
The two mechanisms cover different gaps, as the table shows.
| Failure | DBOS alone | Sume Idempotency-Key |
|---|---|---|
| Process dies between steps | Resumes at the next step | Not needed |
| Process dies after Sume accepted the create, before the step result is stored | Step runs again | Replay returns the original job |
429 or 5xx from Sume | Step retries by policy | Retry adopts the first create |
| Same order started twice | Fixed workflow id returns one run | Same key returns one job |
4xx such as a validation error | should_retry can stop retries | Not applicable: fix the request and send again |
What does the workflow look like?
Each status read is its own short step, so DBOS records every poll and the workflow does not lose its place. Waits use DBOS.sleep in place of time.sleep, so the waiting stays inside DBOS instead of in plain Python. The transient function tells DBOS to stop retrying a 4xx other than 429, which keeps a bad request from burning five attempts.
import os, requests
from dbos import DBOS, DBOSConfig, SetWorkflowID
API = os.environ.get("SUME_API", "https://api.sume.com")
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def get(path: str) -> dict:
return requests.get(f"{API}{path}", headers=H, timeout=30).json()["data"]
def transient(e: BaseException) -> bool:
return not (isinstance(e, requests.HTTPError) and 400 <= e.response.status_code < 500
and e.response.status_code != 429)
@DBOS.step(retries_allowed=True, max_attempts=5, interval_seconds=2, should_retry=transient)
def submit(order_id: str, prompt: str) -> str:
r = requests.post(f"{API}/v1/image-1.0/generate", timeout=30,
json={"prompt": prompt, "mode": "async"},
headers={**H, "Idempotency-Key": f"order-{order_id}-hero-v1"})
r.raise_for_status()
return r.json()["data"]["request_id"]
@DBOS.step(retries_allowed=True, max_attempts=3)
def status(job_id: str) -> dict:
return get(f"/v1/jobs/{job_id}/status")
@DBOS.workflow()
def hero_image(order_id: str, prompt: str) -> str:
job_id = submit(order_id, prompt)
while not (s := status(job_id))["terminal"]:
DBOS.sleep(s.get("next_poll_after_seconds") or 3) # durable sleep
return get(f"/v1/jobs/{job_id}/result")["result"]["artifacts"][0]["url"]
if __name__ == "__main__":
DBOS(config=DBOSConfig(name="sume-hero", system_database_url="sqlite:///dbos.sqlite"))
DBOS.launch()
with SetWorkflowID("order-8823-hero-v1"): # one workflow per order, ever
print(hero_image("8823", "Matte black bottle on marble"))Why is the workflow id set from the order?
SetWorkflowID gives the run a name you control. Starting a workflow with an id that already exists returns that workflow's recorded outcome instead of running again, which is what the second invocation in my test showed: the same URL came back and the server saw no new create request. Combined with an order-derived key, that gives two independent guards for the same property, one on your side and one on Sume's.
Keep the id and the key in step. If you bump the version in the key to get a deliberately new image, bump the workflow id the same way, or DBOS will hand back the old result without ever reaching Sume.
- Derive both the workflow id and the key from the order and a version number.
- Make every status read its own step, and use
DBOS.sleepbetween them. - Honour
next_poll_after_seconds; per-request waits against Sume are short. - Stop retrying permanent
4xxresponses withshould_retry. - Return the artifact URL from the workflow; store bytes elsewhere.
What can still surprise you?
A step that exhausts its attempts raises, and the workflow fails with it unless you catch the exception. Decide whether a failed create should fail the order or park it for a human, and log the Sume request_id from the error body so support can find the call. A 402 is the usual case worth parking: retrying does not change a balance.
Polling is simple, but a long job means a long-lived workflow. If you run many, a webhook that starts the next step is gentler on your own quota, and the DBOS workflow can wait for it. Both paths read the same job object, so the submit step does not change. Whichever you choose, verify the webhook signature before trusting a delivery, and dedupe on the job id, since a delivery can repeat.
Last, test the crash path yourself. Kill the process after the submit step has printed its job id, start it again with the same workflow id, and confirm the job count on your account did not grow. That one experiment shows both guards working, and it is cheap because the idempotent replay creates nothing.
Sources
Related posts
More in Developers
- Dub one Short into 8 languages: Python fan-out and the total cost
Detach and transcribe once, then run one TTS job and one render per language. A Python fan-out and the per-Short bill, from Sume's catalog rates.
- FLUX 3 bounding box to a mask_url: Python region edit on Sume
FLUX 3 Image boxes use [top, left, bottom, right] on a 0-1000 grid. Convert one to an RGBA mask with Pillow and run the region edit on Sume's GPT Image 2.5.
- FLUX 3 Image on OpenRouter: n=1, seed, base64 vs Sume
OpenRouter lists FLUX.3 Image with one image per call, a seed and base64 PNG output. How each differs from Sume's POST /v1/images, where FLUX 3 is not listed.
- FLUX 3 Image on Replicate: safety_tolerance 0-4 vs Sume
Replicate's FLUX 3 Image form has safety_tolerance 0-4, grounding, output_quality and 768sq-4k. Which of those inputs Sume's image API has, and what it returns.
Written by Sume