Postgres SKIP LOCKED job table that submits Sume jobs (Python)

A render_jobs table, a claim query with FOR UPDATE SKIP LOCKED, and an Idempotency-Key per row, so two workers never double-submit a paid Sume job.

4 min readSume
All posts

To run Sume submits from a Postgres table, claim one row at a time with FOR UPDATE SKIP LOCKED, send the row id as the Sume Idempotency-Key, and store the returned job id on the same row. The lock keeps two workers off the same row; the key covers the case the lock cannot, a worker that dies after Sume accepted the job but before the row was updated.

This is the plain-Postgres version of the pattern used by queue extensions. It needs one table, one query and no broker, which suits a team that already has a database and a few thousand generations a day.

Why SKIP LOCKED, and its stated limit

The PostgreSQL documentation describes the clause directly: rows that cannot be locked immediately are skipped, which gives an inconsistent view of the data, so it is not suited to general-purpose reads, but it can be used to avoid lock contention with multiple consumers accessing a queue-like table. A job table is that exact case.

The schema needs a status column with three values you control, a place for the Sume job id, and a claim timestamp so a crashed worker's rows can be put back. The claim query below flips a row to submitting in one statement, so the row lock is held only for that statement and no transaction stays open while the HTTP call runs.

Row states in the job table (read 2026-10-03)
StatusSet byMeaningRecovery
pendingProducer insertWaiting for a workerNone needed
submittingClaim queryA worker owns the row and is calling SumeReset to pending after a stale claimed_at; the same key makes the resubmit safe
submittedWorker, after 202sume_job_id is storedPoll or take a webhook
failedWorker, on a 4xx other than 429The request itself is wrongFix the prompt, then re-insert

The worker

It uses psycopg 3 on an autocommit connection and requests. A 429 or a network error puts the row back to pending and re-raises so the caller can sleep; any other 4xx marks the row failed instead of looping forever. Create the table first: id bigserial primary key, prompt text, status text default 'pending', sume_job_id text, claimed_at timestamptz.

import os, requests, psycopg

HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
CLAIM = """update render_jobs set status='submitting', claimed_at=now()
 where id = (select id from render_jobs where status='pending'
             order by id for update skip locked limit 1)
 returning id, prompt"""

def run_once(conn):
    row = conn.execute(CLAIM).fetchone()
    if not row:
        return False
    job_id, prompt = row
    try:
        r = requests.post("https://api.sume.com/v1/images",
            headers={**HEAD, "Idempotency-Key": f"render-job-{job_id}"},
            json={"model": "sume/auto", "prompt": prompt, "mode": "async"}, timeout=20)
    except requests.RequestException:
        conn.execute("update render_jobs set status='pending' where id=%s", (job_id,))
        raise
    if r.status_code == 202:
        sid = r.json()["data"]["job"]["id"]
        conn.execute("update render_jobs set status='submitted', sume_job_id=%s where id=%s", (sid, job_id))
    elif r.status_code in (429, 503):
        conn.execute("update render_jobs set status='pending' where id=%s", (job_id,))
        raise RuntimeError(f"Sume asked to back off: {r.status_code}")
    else:
        conn.execute("update render_jobs set status='failed' where id=%s", (job_id,))
    return True

# conn = psycopg.connect(os.environ["DATABASE_URL"], autocommit=True)

Pacing the loop

The loop around run_once should respect two Sume limits. Write requests are budgeted per minute by plan, 120 on Free and 300 on Pro per the authentication page, and a full queue answers 429 queue_full. A worker that sleeps on RuntimeError and otherwise loops without delay stays under both; for large backlogs read generation_limits.queue_capacity_remaining from the submit response and stop claiming when it nears zero.

Reaping is one more statement: update render_jobs set status='pending' where status='submitting' and claimed_at < now() - interval '10 minutes'. Because the key is render-job-<id>, a row reaped after the original request actually succeeded returns the original job with idempotency_hit: true and bills nothing extra.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume