Prove a Sume video retry is safe: same key, same job, one charge

A runnable test for a ported Sora worker: submit twice with one Idempotency-Key on /v1/videos, assert the same job id came back, and cancel before it bills.

5 min readSume
All posts

Send the same Idempotency-Key twice on POST /v1/videos and you get the first job back, not a second paid one. The Sume docs state it plainly: a replay returns the original job (Sume video docs, read 2026-10-06). Do not trust that in production on faith. Write the five-line test below and run it once against your own key.

The OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), and a wrapper that retried on timeout without a key is a risk on any replacement. The OpenRouter video wire that Sume copies has no idempotency on this route, so the key is a Sume addition your code has to start using.

What the key changes

Retry behavior with and without a key, Sume docs, read 2026-10-06
SituationWithout a keyWith the same key
Client times out after the server acceptedA retry creates a second job and reserves a second amountThe retry returns the original job
Process crashes and restartsUnknown whether a job existsResubmit the same key to recover the job id
Rate limited (429 rate_limited)A retry may create new workRetry with the same key after backing off
Prompt changes on purposeNew jobUse a new key; the same key with a different body returns 409 conflict

The test

It submits a cheap 3-second Omni clip at 360p twice with one key, asserts the ids match, then cancels. A cancel only works before generation starts, so the job may complete first and bill about $0.12; that is the price of the proof.

import os
import uuid
import requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
key = f"replay-test-{uuid.uuid4()}"
body = {"model": "gemini-omni-flash-1.1", "prompt": "A paper boat on a puddle",
        "aspect_ratio": "9:16", "resolution": "360p", "duration": 3}


def submit() -> str:
    r = requests.post("https://api.sume.com/v1/videos",
                      headers={**H, "Idempotency-Key": key}, json=body, timeout=30)
    r.raise_for_status()
    return r.json()["id"]


first, second = submit(), submit()
assert first == second, (first, second)
print("same job:", first)
c = requests.post(f"https://api.sume.com/v1/jobs/{first}/cancel", headers=H, timeout=30)
print("cancel:", c.status_code)

Choosing the key

  • Derive it from the business event, such as order-8841-hero-clip, so every retry path computes the same string.
  • Never use a timestamp or a random value generated per attempt; that defeats the point.
  • Never reuse a key for a different prompt or model. The same key with a different body is rejected with a 409 conflict, per the Video docs error table.
  • Keep keys per workspace purpose. A fallback chain that tries two models needs one key per model, not one for the chain.

What you still own

The key stops double submission; it does not stop double handling downstream. If a webhook arrives twice, dedupe by job_id before copying a file or notifying a customer. And if a replay comes back as a failed job, do not loop on the same key: change the request, read the error, and decide.

The earlier posts on the retry-on-timeout wrapper and on a fallback chain with one key per model apply the same rule to two common shapes.

Where to run the test

Run the replay test once when you set up an account, and again whenever you change how the key is built. Do not run it in a hot path. It is a check on the behavior you rely on, and the cheapest version of it uses the shortest Omni clip at the lowest resolution, which is the $0.12 job from the earlier post on rounding.

A good place for it is a nightly or pre-release job that uses a separate test key. If the assertion ever fails, treat it as an incident: stop automatic retries until you know why two ids came back for one key.

Remember that the test proves behavior on the day you ran it. Put the date in the test name or the log line, and run it again after a major change on either side, so that a green result always has a date next to it.

  • Use a dedicated key per test run, so a stale key from yesterday cannot fake a pass.
  • Cancel the job after the assertion; the cancel works only before generation starts, so expect that the job may complete and bill instead.
  • Log the first and second ids, and the HTTP status of both submits.
  • If the second call returns an error instead of the first job, read the body and compare it with the docs before you change your code.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume