Hedged requests on a paid video API: same key, one job
Can you hedge a slow Sume submit to cut tail latency? Only with the same Idempotency-Key. Why a second key is a second bill, and what 409 in use means.

Yes, you can hedge a Sume submit, but only if both copies carry the same Idempotency-Key. With one key the worst case is a 409 idempotency_key_in_use on the second copy, which you discard. With two different keys you have asked for two paid generations, and both will be reserved and billed. Hedging is a latency tactic for cheap reads; on a paid create it is a spending decision.
What the key does to the second request
The Format run docs describe the key as a header of up to 255 characters, scoped per Format. A repeat with the same key and the same body returns 200 with idempotency_hit: true and the original run. A repeat with a different body is 409 idempotency_conflict. A repeat that arrives while the first is still being created is 409 idempotency_key_in_use, which is marked retryable and worth waiting about a second on.
That makes the hedge outcome predictable. The first copy to land creates the run; the other copy either replays it or reports that the key is in use. Neither creates a second run. The script below simulates the race with an in-memory store and prints one created job.
| Hedge copy | Result | What to do |
|---|---|---|
| Same key, same body, after creation | 200, idempotency_hit true | Use the replayed run id |
| Same key, still being created | 409 idempotency_key_in_use (retryable) | Wait about 1 s, or just keep the first response |
| Same key, different body | 409 idempotency_conflict | Bug in your code: fix the body, do not retry |
| Different key | Second paid run | Never hedge this way |
import asyncio
async def submit(key, delay, calls):
await asyncio.sleep(delay)
if key in calls:
return 409, "idempotency_key_in_use"
calls[key] = "job_1"
return 202, calls[key]
async def main():
calls = {}
first = asyncio.create_task(submit("order-77-v3", 0.0, calls))
hedge = asyncio.create_task(submit("order-77-v3", 0.05, calls))
print(await first, await hedge, len(calls), "job created")
asyncio.run(main())Why hedging rarely pays on a submit
Hedging helps when the request has a long latency tail and the work is free to repeat. A Sume create returns 202 quickly and the real latency lives in the run, which you poll or receive by webhook. The tail you actually fight is the generation time, and a hedge cannot shorten that: both copies point at the same run.
The key's behaviour after failure matters too. A create that fails with 402 or 503 releases the key, so a hedge that lost to a transient 503 can legitimately win on retry. Under the SDK, createSumeClient retries 408, 429 and 5xx twice with backoff and honours retry-after, but retries a POST only when an Idempotency-Key is present, so the retry you get for free already behaves like a safe hedge.
- Derive the key from your own record (order id plus version), never from a fresh UUID per attempt, so every copy shares it.
- Cap the hedge at one extra copy and delay it past your normal p95 so you do not double the write budget on every call.
- Cancel nothing on the loser: the run belongs to both copies. See the key-in-use post.
A rule to put in code review
Reject any code path that builds the key inside the retry or hedge loop. The key is created once, before the first attempt, and passed down. For long-running agent loops, treat a missing key on a paid create as a bug and fail the call before it leaves the process.
Remember the write budget too: each copy counts against the per-minute write limit for your plan, so a fleet that hedges every submit spends it faster than one that does not.
Where hedging does belong
Hedge reads, not creates. A status read is free of side effects and cheap against a read budget that is 40 times the write budget, so a duplicate read after a slow one is harmless. Even there, keep the hedge small: an extra copy only when the first has outlived your own p95, never as a default on every call. For creates, the better lever is a tighter client timeout plus the Idempotency-Key, so a slow first attempt is simply retried and replayed instead of raced.
Sources
Related posts
More in Developers
- Hey API openapi-ts: generate a Sume client from reference/json
Point @hey-api/openapi-ts at https://api.sume.com/reference/json, pin the version, and send one credential header. When the official SDK is the shorter route.
- Hide the seam in an AI image edit with a feathered composite in Pillow
A hard-edged paste of an AI edit leaves a visible line. Blur the mask a few pixels so the edit fades into the original. Short Pillow script and settings.
- Home Assistant response_variable: read a Sume job id and status
Home Assistant rest_command returns status, content and headers in response_variable. Here is how to pull the Sume job id from it and branch on the status read.
- Home Assistant rest_command 10-second timeout and Sume async jobs
rest_command times out at 10 seconds by default. Sume sync waits cap at 30. Submit with mode async, take the job id, and poll in a second command.
Written by Sume