Find Sume jobs still queued or processing after 30 minutes

A scheduled script lists queued and processing jobs from GET /v1/jobs, picks those older than a cutoff, and prints each so you recover instead of resubmit.

5 min readSume
All posts

List jobs with status=queued and then status=processing, keep those whose created_at is older than your cutoff, and for each print its job id, status and age. A job that has not finished in 30 minutes needs a human or a recovery step, not a resubmit, because the original paid work may still complete.

What to do with a stale job

The list endpoint filters by status and pages with next_cursor. For a stale job, read GET /v1/jobs/{id}/status for next_action and GET /v1/jobs/{id}/events for where it stopped.

Sweep inputs and actions (read 2026-10-04)
InputAction
status=queued, oldCheck queue position and your plan concurrency
status=processing, oldRead events, then keep polling status
cancelable: truePOST /v1/jobs/{id}/cancel is possible
cancelable: falseGeneration started; wait for a terminal state

Script

This script lists and reports only; it never cancels or resubmits. The cutoff is a command-line number of minutes.

import asyncio, json, os, sys, urllib.parse, urllib.request
from datetime import datetime, timedelta, timezone
def fetch(status, cursor):
    q = {"limit": "100", "status": status}
    if cursor: q["starting_after"] = cursor
    req = urllib.request.Request("https://api.sume.com/v1/jobs?" + urllib.parse.urlencode(q),
                                 headers={"x-api-key": os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)["data"]
def stale(minutes):
    cutoff = datetime.now(timezone.utc) - timedelta(minutes=minutes)
    found = []
    for status in ("queued", "processing"):
        cursor = None
        while True:
            page = fetch(status, cursor)
            for j in page["jobs"]:
                made = datetime.fromisoformat(j["created_at"].replace("Z", "+00:00"))
                if made < cutoff:
                    found.append((j["id"], j["status"], round((datetime.now(timezone.utc) - made).total_seconds() / 60)))
            cursor = page.get("next_cursor")
            if not cursor: break
    return found
async def main(minutes):
    rows = await asyncio.to_thread(stale, minutes)
    for job_id, status, age in rows:
        print(f"{job_id} {status} {age} min")
    print(f"{len(rows)} stale jobs")
asyncio.run(main(int(sys.argv[1]) if len(sys.argv) > 1 else 30))

Scheduling

Run it from any scheduler you already have, and send the output to a channel your team watches. The CLI equivalent for a single job is sume jobs watch <job_id>; see recovering after a watch timeout.

Visibility

Jobs created by a different member or key are not visible to your key, so each key's sweep covers only its own jobs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume