Prediction deadlines vs a Sume client deadline: stopping isn't cancel

A client-side deadline stops you watching a Sume job; it does not cancel it. Cancel works only before generation starts, so a missed deadline still bills.

4 min readSume
All posts

On Sume, a deadline in your own code only decides when you stop checking. It does not cancel anything: the job keeps running and keeps billing, and cancel succeeds only before generation starts. Replicate's changelog lists prediction deadlines as a feature (Mar 2026); this post is about the Sume side of the comparison.

The practical question is where a deadline belongs in a client.

Three timeouts that get confused

The docs suggest a client-side deadline of around 20 minutes for video as reasonable, and stress that it is not wait_timeout_seconds, which never exceeds 30 seconds.

Timeouts on Sume jobs (read 2026-10-03)
TimeoutWhat it limitsEffect on the job
wait_timeout_seconds (sync mode)How long the HTTP request blocks; clamped to 0 to 30 sNone. The job continues.
jobs_wait sliceOne MCP wait; default 50 s, cap 55 sNone. Repeat the wait.
Your client deadlineHow long your code keeps pollingNone. The job continues and bills.

What a missed deadline should do

When your deadline passes, you have three honest options: keep waiting, cancel, or hand the job id to something else. Pretending the job ended is not one of them.

  • If the job is queued, try POST /v1/jobs/:id/cancel. It works before generation starts.
  • If it is processing, expect 409 job_generation_already_started with details.cancelable: false.
  • In that case keep the id, report the job as running, and check again later.
  • If the server timed out on your submit, retry with the same Idempotency-Key so you get the original job.

A deadline-aware loop

The official client-subscribe recipe is a loop. Submit with mode: async and an idempotency key, poll GET /v1/jobs/{id}/status, stop on terminal, then read the result.

Add your own deadline outside that loop, and make the exit path return the job id rather than an error that discards it.

import time


def poll_until(get_status, deadline_seconds=1200.0):
    """get_status() returns the parsed status JSON for one job."""
    end = time.monotonic() + deadline_seconds
    delay = 2.0
    while True:
        status = get_status()
        if status.get("terminal"):
            return status, True
        if time.monotonic() >= end:
            return status, False  # still running: keep the job id
        time.sleep(status.get("next_poll_after_seconds") or delay)
        delay = min(delay * 2, 30.0)

What the returned flag is for

When the second value is false, store the job id with its status_url and schedule another check. The job is not lost, and a later pass can read the result once result_ready is true.

A job that finishes after your deadline still finishes, and a later pass can read its result.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume