Test Sume retry code with a unittest fake server: 429 then 202
A stdlib fake server answers 429 with Retry-After, then 202. The test asserts your retry reuses the same Idempotency-Key and waits 7 seconds, with no sleep.

To test Sume retry code without a network or a paid job, start a fake HTTP server on a random local port, queue its replies, and assert on what it saw. The 30-line test below queues a 429 with Retry-After: 7 and then a 202, and it proves two things about the client: the same Idempotency-Key went out twice, and the code asked to wait 7 seconds. The wait is injected as a function, so the test finishes in half a second.
It uses unittest and http.server only. Nothing needs to be installed, and it runs in CI as it is.
What the test must prove
The sample covers the first two rules. The last two are one more reply queue and one more test method each.
| Rule | Assertion |
|---|---|
| Retry a 429 with the same key | The server saw the same Idempotency-Key twice |
| Honor Retry-After | The injected sleep got 7.0 |
| Do not retry a 409 | One request, and the error is raised |
| Stop after a fixed count | The reply queue still holds unused answers |
The test
HTTPServer(('127.0.0.1', 0), Fake) asks the OS for a free port, and srv.server_port tells the test which one. The handler pops one (status, headers) pair from REPLIES for each request and records the key. The thread is a daemon, so a failed test does not hang the run.
import json, threading, unittest, urllib.error, urllib.request
from http.server import BaseHTTPRequestHandler, HTTPServer
REPLIES, KEYS = [], [] # queued (status, headers) and the keys the server saw
class Fake(BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("Content-Length", 0)))
KEYS.append(self.headers.get("Idempotency-Key"))
status, headers = REPLIES.pop(0)
self.send_response(status)
for k, v in headers.items():
self.send_header(k, v)
self.end_headers(); self.wfile.write(b"{}")
def log_message(self, *a): pass
def submit(url, key, sleep):
for _ in range(3):
req = urllib.request.Request(url, data=b"{}", method="POST", headers={"Idempotency-Key": key})
try:
return urllib.request.urlopen(req).status
except urllib.error.HTTPError as e:
if e.code != 429: raise
sleep(float(e.headers.get("Retry-After", 1)))
class RetryTest(unittest.TestCase):
def test_429_then_202_reuses_key_and_waits(self):
srv = HTTPServer(("127.0.0.1", 0), Fake)
threading.Thread(target=srv.serve_forever, daemon=True).start()
REPLIES[:], waits = [(429, {"Retry-After": "7"}), (202, {})], []
self.assertEqual(submit(f"http://127.0.0.1:{srv.server_port}", "k1", waits.append), 202)
self.assertEqual((KEYS, waits), (["k1", "k1"], [7.0]))
srv.shutdown()
unittest.main()Why inject the sleep
A retry loop that calls time.sleep directly makes the test slow and hides the number it was told to wait. Pass the sleep function in, give it waits.append in the test, and read the list afterwards. The same trick works for a clock, if your loop has a deadline.
A fake server is not a copy of Sume. It checks your code against the behavior the docs describe: 429 rate_limited and queue_full both mean retry with the same key, and a replay returns the original job. It cannot prove how the real service behaves under load.
Extending it
- Add a
503reply with no header and assert on your default backoff. - Add a
409reply and assert that the code raised after one request. - Add a reply with a
Retry-AfterHTTP date, and decide what your parser should do.
Running it in CI
Run python3 -m unittest in the folder, or call the file directly. The server binds to 127.0.0.1 on a port the OS picks, so two test runs on the same machine never collide, and nothing is exposed to the network.
The file ends with unittest.main(), which keeps the sample self-contained. In a real project, remove that line and let the test runner discover the class.
Keep one rule for fake servers: they model the contract, not the service. When the contract changes, for example a new error code, update the fake in the same commit as the client.
Add a test for the unhappy path too: a reply queue of four 429 answers should end in an error after your attempt limit, and the server should have seen exactly that many requests.
Sources
Related posts
More in Developers
- The 202 from POST /v1/videos has four fields: where the rest arrives
The Sume submit response returns only id, polling_url, status and model. Usage, unsigned_urls and error come on the poll. A parser that expects no more.
- Threads API video rules: H.264, AAC, edit lists, and what Sume covers
Meta's Threads page lists MP4 or MOV, H.264 or HEVC, AAC, no edit lists and a front moov atom. Which of these Sume documents, and which you must check.
- Threads API video max is 300 s: split a 21-minute talk into parts
The Threads API page lists video up to 300 seconds and 1 GB. Split a 21-minute recording into five trim requests with a small Python script. $0.02 per trim.
- Threads video aspect ratio runs 0.01:1 to 10:1: Sume Timeline sizes
Meta's Threads page allows ratios from 0.01:1 to 10:1 with 9:16 recommended and 1920 px max width. Sizes that Sume Timeline can output inside those bounds.
Written by Sume