Gemini batch create is not idempotent: two jobs vs Sume bulk replay
Gemini's docs say sending the same batch creation request twice creates two batch jobs. Sume bulk runs replay the old queue for the same Idempotency-Key.

Gemini's Batch API does not deduplicate: its docs say that if you send the same creation request twice, two separate batch jobs will be created. A Sume bulk run does: the same Idempotency-Key with the same { concurrency, items } returns 202 and the existing queue, and a different payload under that key returns 409 idempotency_conflict. If a retry loop wraps your batch submit, that difference decides whether a timeout costs you one batch or two.
Gemini's behaviour is from its Batch API page; Sume's is from Bulk runs.
What exactly does Gemini say?
The Batch API page states job creation is not idempotent. Jobs move through JOB_STATE_PENDING, JOB_STATE_RUNNING, JOB_STATE_SUCCEEDED, JOB_STATE_FAILED, JOB_STATE_CANCELLED and JOB_STATE_EXPIRED, the last after more than 48 hours pending or running. Results are kept for download for 6 weeks, and the target turnaround is 24 hours. Nothing on the page we read describes a client-supplied key that makes creation safe to retry, so if you retry a create that timed out, check your jobs list before you resend.
How does Sume handle the same retry?
Send Idempotency-Key on every create. For bulk runs the key is scoped to one Format, and the body spelling idempotency_key also works, with the header winning if you send both.
| Replay | Result |
|---|---|
| Same key, same concurrency and items | 202 and the existing queue |
| Same key, different payload | 409 idempotency_conflict, details.queue_id names the original |
| New key, same payload | A second queue, and a second set of runs |
Why does the bulk replay stay at 202?
Unlike a single run, where an idempotent replay is 200 with idempotency_hit: true, a bulk replay stays 202, and the queue object has no idempotency_hit field. So to detect a replay, compare the returned id with the queue id you stored for that key. The docs also warn that replaying a spent key returns the old queue, so mint a fresh key per batch you really intend as a new batch.
What key should I derive?
Derive the key from the thing being made, not from a random value per attempt. A hash of the batch contents means a retry reuses the key, and an edited batch gets a new one.
import hashlib, json, os, requests
items = [{"instruction": f"clip {i}"} for i in range(1, 6)]
body = {"concurrency": 2, "items": items}
key = hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()[:32]
url = "https://api.sume.com/v1/formats/chase/product-promo/bulk-runs"
for attempt in range(3):
try:
r = requests.post(url, json=body, timeout=30, headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": key})
break
except requests.Timeout:
continue
else:
raise SystemExit("no response, check the queue before resending")
print(r.status_code, r.json()["data"]["id"])Because the key is a content hash, an unchanged retry returns the first queue; an edited item list gets a new key and a new queue, so the 409 never fires. If you want edits to be refused instead, derive the key from a batch id of your own. For a failed single run, the docs say to retry with a new key rather than the old one, because the old key is bound to the receipt you already have.
Sources
Related posts
More in Comparisons
- Flow charges 7 to 15 credits for Omni Flash; Sume bills dollars
Google Flow prices Gemini Omni Flash 720p at 7 to 15 credits by length. Sume bills the same model per second in dollars; clip costs from 3 to 10 seconds.
- Flow video edit costs 40 credits a try; Sume bills per second
Flow charges 40 credits per Gemini Omni video edit. Sume's video_url edit bills output seconds: $0.625 for 5 seconds, $1.25 for 10 at 720p.
- GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs
GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.
- gpt-image-2.5 partial image streaming vs Sume's 400
OpenAI streams 0 to 3 partial images for gpt-image-2.5. Sume returns 400 streaming_not_supported for stream: true. What to send instead: a job and a poll.
Written by Sume