Record the Idempotency-Key before you POST a Sume video job

A worker that dies after submitting a video job loses the job id. Write the key to a ledger first, then POST; a retry returns the same job. Tested SQLite code.

5 min readSume
All posts

To avoid paying twice when a worker dies mid-submit, write the Idempotency-Key to a ledger row before you POST the video job, then send the request with that exact key, then store the returned job id on the same row. If the process crashes between the POST and the update, the restart finds a ledger row with a key and no job id, repeats the POST with the same key, and gets the original job back. This is the contract in Sume jobs and results: a retry with the same key and the same payload returns the original job, and a different payload under that key returns 409 idempotency_conflict.

The order matters because the dangerous window is small but real. A video job spends credits, the 202 response carries the only copy of the job id you have, and a timeout, an OOM kill or a deploy can cut the connection after Sume accepted the work.

What does the ledger hold?

The key is not a random value you invent per attempt. It is a function of the work: the order, the shot and the version of the prompt. If the key changes on retry, the server cannot recognise the repeat, and the protection is gone. If the payload changes while the key stays, you get the 409, which is correct behaviour and a signal to mint a new key for the new intent.

Persist three columns: key (primary key), payload_hash, and job_id (nullable). The hash lets a restart detect that the code now builds a different payload for the same key before the server has to tell you.

How does the code look?

import sqlite3, hashlib, json

db = sqlite3.connect(":memory:")
db.execute("create table ledger(key text primary key, hash text, job_id text)")

def submit(key, payload, post):
    h = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    db.execute("insert or ignore into ledger(key, hash) values (?, ?)", (key, h))
    q = "select hash, job_id from ledger where key=?"
    stored_hash, job_id = db.execute(q, (key,)).fetchone()
    if stored_hash != h:
        raise ValueError("payload changed under an existing key: mint a new key")
    if job_id:
        return job_id
    job_id = post(key, payload)
    db.execute("update ledger set job_id=? where key=?", (job_id, key))
    db.commit()
    return job_id
server = {}
def fake_post(key, payload):
    return server.setdefault(key, "job_%d" % (len(server) + 1))
def crashing_post(key, payload):
    fake_post(key, payload)
    raise TimeoutError("dropped after the server accepted the job")

body = {"model": "seedance-2.5", "prompt": "a kettle", "duration": 8}
try: submit("order-7:shot-1:v1", body, crashing_post)
except TimeoutError: pass
print(submit("order-7:shot-1:v1", body, fake_post), "jobs created:", len(server))
assert len(server) == 1

Why test with a crash in the middle?

The fake server in the test behaves the way Sume documents: a repeated key returns the first job. Run the same shape against the real API in a staging key to confirm, using a cheap clip, and you have a regression test for the exact incident that otherwise only shows up as a surprise line on an invoice.

Do not delete ledger rows when a job finishes. A completed row is what answers a late retry from a queue that redelivers messages, and it is the record you search when someone asks which request produced a clip.

Where a crash can land, and what the ledger does, from Sume jobs and results (read 2026-10-06)
Crash pointLedger stateOn restart
Before the ledger insertNo rowTreat as new work
After insert, before POSTKey, no job idPOST with the same key
After POST, before updateKey, no job idSame POST returns the original job
After updateKey and job idSkip the POST, poll the job

What the server does with the key

The /v1/videos route supports the Idempotency-Key header, and the docs list a 409 for a key reused with a different body. That is the same pair of outcomes the ledger relies on: an identical body replays the original job, and a changed body is refused (Sume docs: Errors and credits, read 2026-10-06). A job that was rejected at submit never reserved credits, so a ledger row with a key and no job id is safe to retry as is.

The jobs guidance adds one rule that the ledger encodes for you: if you lose track of a submit, retry the submit with the same key; do not submit a new paid job for the same intent. After the job id is known, switch to polling GET /v1/jobs/{id}/status until it is terminal, and read the result only when result_ready is true. Store the terminal status on the ledger row too, so a restart does not poll a job that already finished.

What does the ledger not cover?

Two limits to keep in mind. The ledger prevents duplicate submits, not duplicate downloads or duplicate publishes; those need their own dedupe on the job id, which a webhook handler or a poller can supply. And a key is scoped to the payload you sent: if a product change means the prompt text differs on purpose, include a version number in the key so the new intent gets a new key instead of a conflict.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume