Resume an Omni batch after a crash with stable idempotency keys

A worker dies halfway through 40 Omni clips. Build Idempotency-Key from the item id so a rerun resubmits safely, and treat 409 and 429 as signals, not failures.

5 min readSume
All posts

To resume a crashed Omni batch on Sume, give each item an Idempotency-Key built from your own stable item id, never a random value, and resubmit the whole list. Sume documents retrying a submit with the same key after queue_full, and returns 409 idempotency_conflict when a key is reused with a different payload. So a rerun is safe if each item's body is unchanged, and a 409 tells you the item changed.

What the three signals mean

Generation admission lists them. 429 queue_full means you already hold the maximum accepted jobs, so wait and retry with the same key. 402 insufficient_credits means the reservation failed, so stop. 409 idempotency_conflict means this key was used with a different body. Retry rules are in Errors and credits.

The docs do not state how long a key is remembered, so do not rely on a rerun days later; keep your own ledger of job ids.

Rerun outcomes per item when a batch restarts (Sume docs, read 2026-10-04)
Situation on rerunResponseWhat your code does
Same key, same body, already acceptedRetry of the same requestRecord the returned job id
Same key, edited prompt409 idempotency_conflictGive the edited item a new key suffix
Queue full429 queue_fullSleep, poll held jobs, retry same key
Balance empty402 insufficient_creditsStop the batch

A rerunnable submit

The key is the batch name plus the item id plus a version number you bump only when you edit the prompt. The loop stops on 402 and sleeps on 429.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
ITEMS = {"intro": "A lighthouse at dawn", "outro": "Waves on a rocky shore"}

def submit(item_id, prompt, version=1):
    body = {"model": "gemini-omni-flash-1.1", "prompt": prompt,
            "resolution": "360p", "duration": 3, "aspect_ratio": "9:16"}
    key = f"spring-batch-{item_id}-v{version}"
    while True:
        r = requests.post(f"{API}/v1/video-router/generate", json=body,
                          headers={**H, "Idempotency-Key": key}, timeout=30)
        if r.status_code == 429:
            time.sleep(int(r.headers.get("retry-after", 10)))
            continue
        if r.status_code in (402, 409):
            raise SystemExit(f"{item_id}: stop, HTTP {r.status_code}")
        r.raise_for_status()
        return r.json()["data"]["job"]["id"]

ledger = {name: submit(name, p) for name, p in ITEMS.items()}
print(ledger)

After the rerun

Write each returned job id to disk the moment you have it, then poll them using the interval in Jobs and results. A completed Omni clip is billed at provider list times 1.25 per output second per the Video Router, so a double submit is the cost you are guarding against.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume