One Idempotency-Key for an Omni 360p draft and 720p final: it 409s

Reusing the draft's Idempotency-Key on the 720p final returns 409 idempotency_conflict. Key naming that keeps draft and final jobs apart, with a Python helper.

4 min readSume
All posts

No: if you send the same Idempotency-Key on a 360p draft and then on the 720p final, Sume does not make a second job. The two bodies differ (resolution at least), so the second submit fails with 409 idempotency_conflict. A key identifies one operation. Put the tier in the key and the pair stops colliding.

Google's 360p draft mode for Gemini Omni 1.1 Flash is described as up to 60% faster and a third of the cost of the standard 720p mode, and it is meant for rapid prototyping and storyboard iteration (Google, read 2026-10-05). A draft-then-final loop therefore submits the same prompt twice with a different resolution. That is exactly the pattern that trips the key rule.

What Sume does with the key

The Video generation docs say to send Idempotency-Key to make retries safe, and that a replay returns the original job. The Generation admission docs list the failure case: a client that uses the same key again for a different operation or payload gets 409 idempotency_conflict, and the fix is to reuse a key only for an exact retry.

So there are two outcomes for a reused key. Same body: you get the original job back, no new charge. Different body: you get the 409 and no job. A 409 is the safe failure, but if your code treats it as a generic error it may retry forever or, worse, treat the draft job as the final.

A key scheme that works

Do not hash only the prompt. Two tiers of the same prompt would share a key. Hash the whole request body if you want it generated, because the key rule is about the whole payload.

  • Include the shot id: lisbon-s03.
  • Include the tier: draft or final.
  • Include a version you bump when the prompt changes: v2.
  • Result: lisbon-s03-draft-v2 and lisbon-s03-final-v2. A network retry reuses the exact key and body. A new prompt bumps the version. A promoted draft changes the tier.

Python helper

The request goes to POST /v1/videos with model: gemini-omni-flash-1.1. The model takes 3 to 10 seconds, 360p to 4K, in 16:9 or 9:16 (Sume docs: Video Router, read 2026-10-05).

import os, requests

URL = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(shot, tier, version, prompt, resolution):
    body = {
        "model": "gemini-omni-flash-1.1",
        "prompt": prompt,
        "resolution": resolution,
        "duration": 10,
        "aspect_ratio": "9:16",
    }
    key = f"{shot}-{tier}-v{version}"
    r = requests.post(URL, json=body, headers={**AUTH, "Idempotency-Key": key})
    if r.status_code == 409:
        raise RuntimeError(f"key {key} was used for a different body")
    r.raise_for_status()
    return r.json()["id"]

prompt = "Slow push-in on a ceramic mug, steam rising, morning light"
draft = submit("mug-s01", "draft", 1, prompt, "360p")
final = submit("mug-s01", "final", 1, prompt, "720p")
print(draft, final)

Two cautions

A different key means a new paid job. Sume reserves the estimated amount at submit and refunds a failed job, but a second key for the same shot is a second charge if the first succeeded. Log the key next to the job id so a retry after a timeout reuses it.

Also keep the draft and final in your own table. The final is a new generation, not an upscale of the draft, and the docs say no v1 model accepts seed, so the 720p result will not be the same clip you approved. That limit is why drafts check the idea, not the exact frames. See the 360p draft cost breakdown for what the draft pass costs.

Retries after a timeout

The key earns its keep on the retry path. Say your client times out after the POST for the 360p draft was sent. You do not know whether Sume accepted it. Resend with the same key and the same body: if the job exists, you get it back, and if it does not, you get a new one. In neither case do you pay twice. That only works if the body is the same operation, so build the body in one function and call it for both the first try and the retry, not by editing a shared dict in place between calls.

The Jobs and results docs make the same point in one line: use the same key again only for the same operation and payload. For a batch, derive the key from the shot id and tier instead of a random UUID, because a random UUID generated inside the retry loop is a new key every time and defeats the check.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume