One Idempotency-Key for an Omni 360p draft and 720p final: it 409s
Reusing the draft's Idempotency-Key on the 720p final returns 409 idempotency_conflict. Key naming that keeps draft and final jobs apart, with a Python helper.

No: if you send the same Idempotency-Key on a 360p draft and then on the 720p final, Sume does not make a second job. The two bodies differ (resolution at least), so the second submit fails with 409 idempotency_conflict. A key identifies one operation. Put the tier in the key and the pair stops colliding.
Google's 360p draft mode for Gemini Omni 1.1 Flash is described as up to 60% faster and a third of the cost of the standard 720p mode, and it is meant for rapid prototyping and storyboard iteration (Google, read 2026-10-05). A draft-then-final loop therefore submits the same prompt twice with a different resolution. That is exactly the pattern that trips the key rule.
What Sume does with the key
The Video generation docs say to send Idempotency-Key to make retries safe, and that a replay returns the original job. The Generation admission docs list the failure case: a client that uses the same key again for a different operation or payload gets 409 idempotency_conflict, and the fix is to reuse a key only for an exact retry.
So there are two outcomes for a reused key. Same body: you get the original job back, no new charge. Different body: you get the 409 and no job. A 409 is the safe failure, but if your code treats it as a generic error it may retry forever or, worse, treat the draft job as the final.
A key scheme that works
Do not hash only the prompt. Two tiers of the same prompt would share a key. Hash the whole request body if you want it generated, because the key rule is about the whole payload.
- Include the shot id:
lisbon-s03. - Include the tier:
draftorfinal. - Include a version you bump when the prompt changes:
v2. - Result:
lisbon-s03-draft-v2andlisbon-s03-final-v2. A network retry reuses the exact key and body. A new prompt bumps the version. A promoted draft changes the tier.
Python helper
The request goes to POST /v1/videos with model: gemini-omni-flash-1.1. The model takes 3 to 10 seconds, 360p to 4K, in 16:9 or 9:16 (Sume docs: Video Router, read 2026-10-05).
import os, requests
URL = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def submit(shot, tier, version, prompt, resolution):
body = {
"model": "gemini-omni-flash-1.1",
"prompt": prompt,
"resolution": resolution,
"duration": 10,
"aspect_ratio": "9:16",
}
key = f"{shot}-{tier}-v{version}"
r = requests.post(URL, json=body, headers={**AUTH, "Idempotency-Key": key})
if r.status_code == 409:
raise RuntimeError(f"key {key} was used for a different body")
r.raise_for_status()
return r.json()["id"]
prompt = "Slow push-in on a ceramic mug, steam rising, morning light"
draft = submit("mug-s01", "draft", 1, prompt, "360p")
final = submit("mug-s01", "final", 1, prompt, "720p")
print(draft, final)Two cautions
A different key means a new paid job. Sume reserves the estimated amount at submit and refunds a failed job, but a second key for the same shot is a second charge if the first succeeded. Log the key next to the job id so a retry after a timeout reuses it.
Also keep the draft and final in your own table. The final is a new generation, not an upscale of the draft, and the docs say no v1 model accepts seed, so the 720p result will not be the same clip you approved. That limit is why drafts check the idea, not the exact frames. See the 360p draft cost breakdown for what the draft pass costs.
Retries after a timeout
The key earns its keep on the retry path. Say your client times out after the POST for the 360p draft was sent. You do not know whether Sume accepted it. Resend with the same key and the same body: if the job exists, you get it back, and if it does not, you get a new one. In neither case do you pay twice. That only works if the body is the same operation, so build the body in one function and call it for both the first try and the retry, not by editing a shared dict in place between calls.
The Jobs and results docs make the same point in one line: use the same key again only for the same operation and payload. For a batch, derive the key from the shot id and tier instead of a random UUID, because a random UUID generated inside the retry loop is a new key every time and defeats the check.
Sources
Related posts
More in Developers
- Save a 30-second Wan 3.0 clip to Cloudflare R2 from Node
Fetch a finished Wan 3.0 render from Sume and write it to Cloudflare R2 with the AWS S3 client. Shows the R2 endpoint config and the 30 s price at three tiers.
- script_run for one TTS clip per sentence: limits to set first
Fan out one tts_create per sentence inside Sume script_run, with timeout_seconds, max_calls and max_paid_calls set, then wait on the child jobs with jobs_wait.
- script_run error script_tool_forbidden: discovery calls belong outside
script_tool_forbidden means a Sume script called a discovery tool (tools_list, tools_schema, mcp_health, search_tools) or script_run. Call them from the turn.
- script_text, words, cues or segments: which caption input to send
Sume captions take only one of script_text, words, cues, segments. Your pick decides whether speech-to-text runs and what happens on a silent clip.
Written by Sume