Prove a Sume retry makes one job: drop the response on purpose
A 28-line fake server creates the job, then hangs up before replying. A stable Idempotency-Key leaves one job; a fresh key per retry leaves two. Runs anywhere.

Build a fake server that creates the job, then closes the connection without answering, and run your submit code against it. If your code reuses one Idempotency-Key, the server ends with one job; if it makes a new key per attempt, the same fault leaves two. The Sume docs promise that a retry with the same key returns the original job instead of billing a second one, and this test checks that your client keeps its half of the deal.
This is the failure that costs money in production: the network drops the response, not the request.
Why this fault matters
The Jobs and results page tells clients that a retry after a client-side timeout or network failure must send the same Idempotency-Key, and that you must not submit a new paid job for the same intent. A client cannot tell a request that never arrived from a request whose answer got lost. Only the key makes the two cases safe to treat alike.
The job record also echoes the key as idempotency_key, described in the OpenAPI file as the join column for recovery, because list pages are newest-first and not in your submit order.
The runnable fault harness
The server stores jobs by key, creates the job first, and swallows one response per DROPS count. The client takes a function that makes the key, so the same harness can test both behaviours. Python's server logs each request to stderr, which is noise only.
import json, threading, urllib.request, uuid
from http.server import BaseHTTPRequestHandler, HTTPServer
JOBS, DROPS = {}, [1] # DROPS: how many responses to swallow after accepting
class Fake(BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers["content-length"]))
job = JOBS.setdefault(self.headers["idempotency-key"], f"job_{len(JOBS) + 1}")
if DROPS[0]:
DROPS[0] -= 1
return self.connection.close() # job created, client never hears
body = json.dumps({"data": {"job_id": job}}).encode()
self.send_response(202)
self.send_header("content-length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def submit(url, make_key, tries=3):
for _ in range(tries):
req = urllib.request.Request(url, b"{}", {"Idempotency-Key": make_key()})
try:
return json.load(urllib.request.urlopen(req, timeout=5))["data"]["job_id"]
except OSError:
pass # dropped connection: retry
srv = HTTPServer(("127.0.0.1", 0), Fake)
threading.Thread(target=srv.serve_forever, daemon=True).start()
url = f"http://127.0.0.1:{srv.server_port}/"
print("stable key:", submit(url, lambda: "intent-1"), "jobs:", len(JOBS))
DROPS[0] = 1
print("fresh key:", submit(url, lambda: str(uuid.uuid4())), "jobs:", len(JOBS))Reading the result
- The stable-key row is the pass condition for your own client. Replace
submitwith your real function and keep the fake server. - The fresh-key row shows what a retry wrapper does when it builds the key inside the loop, which many generic retry libraries do through a request hook.
- Add a variant that drops the response three times in a row to test your retry limit and your final error path.
| Client behaviour | Attempts | Jobs created |
|---|---|---|
| Same key on every try | 2 | 1 |
| New uuid4 on every try | 2 | 2 |
Turn the fault into a regression test
Wrap the harness in your test runner, assert len(JOBS) == 1, and run it in CI with no network and no credits. That is the cheapest guard against the slow regression where someone swaps the HTTP client for one with its own retry hook. Pair it with a real run on the development host now and then to confirm the server side still returns the original job for a repeated key.
If your client crashes between submit and storing the id, you can recover from the other side: list jobs and match your own key against idempotency_key.
Sources
Related posts
More in Developers
- Fit an AI voiceover to a target length with Sume TTS speed 0.6 to 1.5
To hit a target duration, generate once, measure the audio, then set generation_config.speed to measured over target, clamped to Sume's 0.6 to 1.5 range.
- A for await loop over Sume job status: one async generator, 3 runtimes
Wrap Sume status polling in an async generator and read it with for await. One fetch-only file ran unchanged on Node, Bun and Deno and honors the poll hint.
- Generate a Go client from the Sume OpenAPI with oapi-codegen
The full Sume spec trips oapi-codegen on a [number, null] type, but filtering to the operations you use works. Config, generated names and a short caller.
- GitHub Actions concurrency for a Sume render job: the cancel trap
cancel-in-progress stops your workflow, not the Sume job it submitted. Group by branch, cancel via POST /v1/jobs/{id}/cancel in a final step, and cap the run.
Written by Sume