A sensor for a Sume Format run: poll, back off, time out

Poll a Sume Format run with a doubling backoff up to 60 seconds and a hard deadline. Here is the loop, the terminal states, and a 45 minute limit.

4 min readSume
All posts

To wait on a Sume Format run from a workflow tool, poll the receipt's status_url with a delay that doubles up to 60 seconds, stop on a terminal status, and set a hard deadline of your own. For long-form video, 45 minutes is a sensible limit, since runs typically take 15 to 30 minutes.

A webhook is the better primary signal when your workflow can receive one. A sensor is the right tool when it cannot, when you want a hard timeout in the same place as the task, or as the safety net behind a webhook.

The same logic fits any scheduler that has a sensor or a wait step. Only the wrapper changes; the three decisions below do not.

The states you are waiting for

A run is queued, processing, then ends in completed, failed, canceled, or skipped. The receipt also carries next_action, which is poll_status, retry_later, or none, and a queue block whose state is waiting or runtime_unavailable. The queue position is always null, so do not build a progress bar on it. The Format runs page defines all of these.

events_url returns a polled phase timeline of preparing, running, and finalizing. It is not a stream; there is no server-sent events feed or log tail, so a sensor that wants more detail polls it too, but a status poll is all you need to decide.

The loop

The code below uses only the Python standard library. It reads the key from SUME_API_KEY, backs off from 2 seconds, and raises when your deadline passes. It returns the final receipt so the next step can read result_url and usage.

import json, os, time, urllib.request

TERMINAL = {"completed", "failed", "canceled", "skipped"}

def wait_for_run(status_url: str, deadline_s: float = 2700.0) -> dict:
    key = os.environ["SUME_API_KEY"]
    start, delay = time.monotonic(), 2.0
    while True:
        req = urllib.request.Request(status_url, headers={"Authorization": f"Bearer {key}"})
        with urllib.request.urlopen(req, timeout=30) as resp:
            receipt = json.load(resp)
        if receipt.get("status") in TERMINAL:
            return receipt
        if time.monotonic() - start > deadline_s:
            raise TimeoutError("run still going, cancel it via cancel_url")
        time.sleep(delay)
        delay = min(delay * 2, 60.0)

Choosing the deadline

Sume force-finalizes a run at expires_at, which is 90 minutes after creation, or sooner when a run older than 25 minutes has been silent for 10. Your own deadline should be shorter than that, because a sensor that waits for the platform to give up has already spent the time. Forty-five minutes covers the typical 15 to 30 minute video with room to spare.

On a timeout, call the cancel URL and record the outcome. Cancel is idempotent and returns a cancel_effect of canceled or no_op. You pay for generation completed before the cancel.

Return the receipt, not just the status. Downstream steps need result_url, usage, and the failure code, and fetching them again costs reads you did not need to spend.

Sensor settings for video runs, from the Sume docs (read 2026-10-10)
SettingValueSource
First delayA few secondsYour choice
BackoffDoubling, capped at 60 secondsRuns page
Typical long-form time15 to 30 minutesRuns page
Your hard deadline45 minutesYour choice, under the 90 minute expiry
On timeoutCancel, then recordCancel URL

Read budget

Polling is cheap. One 30 minute run polled this way makes about 35 requests: six in the first minute as the delay doubles, then roughly one a minute. Reads get 40 times the write budget, so even the Free plan's 4800 reads a minute is far beyond what a sensor needs. The limit you will feel first is the create budget, which is 120 writes a minute on Free.

Test the timeout path on purpose. A sensor that has never timed out in staging will surprise you in production, so point it at a deliberately tiny deadline once.

After the sensor

Check the status before reading the result. A completed run has a result_url; a failed run has a failure code and a usage block. Hand the failed case to the dead-letter table, not to a blind retry.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume