Prove a Sume video retry is safe: same key, same job, one charge
A runnable test for a ported Sora worker: submit twice with one Idempotency-Key on /v1/videos, assert the same job id came back, and cancel before it bills.

Send the same Idempotency-Key twice on POST /v1/videos and you get the first job back, not a second paid one. The Sume docs state it plainly: a replay returns the original job (Sume video docs, read 2026-10-06). Do not trust that in production on faith. Write the five-line test below and run it once against your own key.
The OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), and a wrapper that retried on timeout without a key is a risk on any replacement. The OpenRouter video wire that Sume copies has no idempotency on this route, so the key is a Sume addition your code has to start using.
What the key changes
| Situation | Without a key | With the same key |
|---|---|---|
| Client times out after the server accepted | A retry creates a second job and reserves a second amount | The retry returns the original job |
| Process crashes and restarts | Unknown whether a job exists | Resubmit the same key to recover the job id |
| Rate limited (429 rate_limited) | A retry may create new work | Retry with the same key after backing off |
| Prompt changes on purpose | New job | Use a new key; the same key with a different body returns 409 conflict |
The test
It submits a cheap 3-second Omni clip at 360p twice with one key, asserts the ids match, then cancels. A cancel only works before generation starts, so the job may complete first and bill about $0.12; that is the price of the proof.
import os
import uuid
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
key = f"replay-test-{uuid.uuid4()}"
body = {"model": "gemini-omni-flash-1.1", "prompt": "A paper boat on a puddle",
"aspect_ratio": "9:16", "resolution": "360p", "duration": 3}
def submit() -> str:
r = requests.post("https://api.sume.com/v1/videos",
headers={**H, "Idempotency-Key": key}, json=body, timeout=30)
r.raise_for_status()
return r.json()["id"]
first, second = submit(), submit()
assert first == second, (first, second)
print("same job:", first)
c = requests.post(f"https://api.sume.com/v1/jobs/{first}/cancel", headers=H, timeout=30)
print("cancel:", c.status_code)Choosing the key
- Derive it from the business event, such as
order-8841-hero-clip, so every retry path computes the same string. - Never use a timestamp or a random value generated per attempt; that defeats the point.
- Never reuse a key for a different prompt or model. The same key with a different body is rejected with a 409 conflict, per the Video docs error table.
- Keep keys per workspace purpose. A fallback chain that tries two models needs one key per model, not one for the chain.
What you still own
The key stops double submission; it does not stop double handling downstream. If a webhook arrives twice, dedupe by job_id before copying a file or notifying a customer. And if a replay comes back as a failed job, do not loop on the same key: change the request, read the error, and decide.
The earlier posts on the retry-on-timeout wrapper and on a fallback chain with one key per model apply the same rule to two common shapes.
Where to run the test
Run the replay test once when you set up an account, and again whenever you change how the key is built. Do not run it in a hot path. It is a check on the behavior you rely on, and the cheapest version of it uses the shortest Omni clip at the lowest resolution, which is the $0.12 job from the earlier post on rounding.
A good place for it is a nightly or pre-release job that uses a separate test key. If the assertion ever fails, treat it as an incident: stop automatic retries until you know why two ids came back for one key.
Remember that the test proves behavior on the day you ran it. Put the date in the test name or the log line, and run it again after a major change on either side, so that a green result always has a date next to it.
- Use a dedicated key per test run, so a stale key from yesterday cannot fake a pass.
- Cancel the job after the assertion; the cancel works only before generation starts, so expect that the job may complete and bill instead.
- Log the first and second ids, and the HTTP status of both submits.
- If the second call returns an error instead of the first job, read the body and compare it with the docs before you change your code.
Sources
Related posts
More in Developers
- Sume job statuses: three vocabularies and a Python normalizer
Ported Sora code checks one status word. Sume has pending, queued and IN_QUEUE depending on the route. A table, and a normalizer you can run with no network.
- Timeout for AI video jobs: set a deadline in your worker, not HTTP
A Sora-era HTTP timeout of 10 minutes breaks on Sume: sync waits cap at 30 s while jobs run for minutes. Use a client deadline and a poll. Python example.
- Callback or polling for a ported video worker: pick by job count
Replacing a Sora polling worker on Sume: when callback_url beats polling, when polling is enough, and one function that does both.
- Split a long video into Shorts episodes: Python trim ranges
Generate back-to-back video trim ranges for a long vertical video with a tail rule, so no episode is a sliver. Respects the 1800 s source and 900 s output caps.
Written by Sume