Restate durable handler: submit a Sume job once, then poll safely

Restate stores completed steps and resumes after failure. Make the Sume submit one stored step, then poll the returned job id without a second charge.

5 min readSume
All posts

The answer

Restate's welcome page lists durable execution as its first capability: code automatically stores completed steps and resumes from where it left off when recovering from failures. So the pattern for a Sume generation is to make the submit one stored step that returns the job id, then poll or wait on that id in later steps.

Anything stored survives a restart, so the paid submit is not repeated for a step that already completed. For the narrow case where a failure hits mid-request, pair it with Sume's Idempotency-Key header.

Why the job id is the right thing to store

The same page also lists reliable communication with exactly-once semantics for calls between services. That is about Restate's own invocations. Your call to Sume is an outside HTTP request, so it only inherits that guarantee if you add your own. Storing the returned job id gives you a durable reference to work that continues on Sume's side regardless of what happens to your process.

Sume's job model fits that. A submit with mode: async returns a status_url, a result_url and a next_action of poll_status; the job keeps running whether or not your handler is alive.

Restate capabilities mapped to a Sume job (read 2026-10-03)
Restate capabilityUse with Sume
Durable executionJournal the submit step and its job id
Built-in stateKeep the order to job id mapping
Reliable communicationYour own handlers; add Idempotency-Key for the outside call
Resume after failureRe-read status from the saved id, never resubmit

Handler outline

Use three named steps in the handler. The first builds the key from your order and revision and submits. The second reads GET /v1/jobs/{id}/status and branches on terminal; when it is false, wait at least next_poll_after_seconds before the next read. The third calls /result once result_ready is true.

Return the result URL or the stored asset reference from the handler, not the raw status payload. Status bodies change as the job moves; a result is what the caller actually wants.

Failure branches to handle

A terminal job can be completed, failed or canceled. The status response sets next_action to inspect_events for the last two, which points you at the job's events list for the reason. Treat 402 insufficient_credits as a business error to surface, not something to retry, and treat 429 as retryable with the delay from retry-after.

A queue_full 429 is a capacity signal rather than a rate limit, so a short fixed backoff will not help; the errors page explains the difference.

Keep the wait cheap

Prefer a webhook when your service is reachable: Sume sends signed terminal events only, so one delivery replaces many polls, and the handler can still fall back to a status read if a delivery is late. Because delivery is a convenience and not the only source of truth, always make the status read the arbiter.

A rule for retries inside the handler

Let the runtime retry the whole step, not a loop you write inside it. If you retry by hand inside the step and the step is itself retried, attempts multiply. Keep the step to a single HTTP request that either returns a job id or throws, and let one layer own the retry policy.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume