Test a Sume poll loop against a fake local server in Python
Test your Sume job poll loop with a stdlib http.server that answers queued, processing, completed. No network, no credits and no real waits. Runnable code.

You can test a Sume poll loop without spending credits by pointing it at a local http.server that returns the same shape: a job that is queued, then processing, then completed with terminal set to true. The loop stops on terminal and injects a no-op sleep, so the test takes milliseconds.
This is a stand-in, not a recording of the real API. It copies only fields the docs describe on GET /v1/jobs/{id}: sume_status, terminal and result_ready.
The fake and the loop
The server walks through three states. The poll function takes a base URL and a sleep function so tests do not wait. It printed ['queued', 'processing', 'completed'].
The server binds to port 0 so the operating system picks a free port, which keeps tests from colliding when they run in parallel. The loop takes its sleep as an argument, so production passes a real sleep and tests pass a function that does nothing; the same trick lets you assert how long the loop would have waited. Because the handler uses a shared iterator, each test should build its own server and its own sequence. When the sequence is exhausted the sample keeps answering completed, which is a simple way to avoid an exception in the handler thread.
import http.server, itertools, json, threading, urllib.request
STEPS = itertools.chain(["queued", "processing"], itertools.repeat("completed"))
class Fake(http.server.BaseHTTPRequestHandler):
def do_GET(self):
status = next(STEPS)
body = json.dumps({"sume_status": status, "terminal": status == "completed",
"result_ready": status == "completed"}).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(body)
def log_message(self, *args): pass
def poll(base: str, sleep=lambda s: None) -> list[str]:
seen = []
while True:
with urllib.request.urlopen(f"{base}/v1/jobs/job_1") as r:
data = json.load(r)
seen.append(data["sume_status"])
if data["terminal"]:
return seen
sleep(2)
srv = http.server.HTTPServer(("127.0.0.1", 0), Fake)
threading.Thread(target=srv.serve_forever, daemon=True).start()
print(poll(f"http://127.0.0.1:{srv.server_port}")) # ['queued', 'processing', 'completed']
srv.shutdown()What to add in real tests
Extend the fake with the failure cases you care about.
| Fake response | Loop should |
|---|---|
| Status 503 or a 524 with HTML | Retry the same job id |
| sume_status failed, terminal true | Stop and read the job error |
| sume_status canceled, terminal true | Stop; do not fetch a result |
| next_poll_after_seconds 30 | Sleep that long, not your default |
Why a fake beats mocking the client
Using a real socket exercises timeouts, status handling and body parsing, which mocks skip. It also catches the common bug of looping forever because you waited for completed and a job ended failed or canceled. Sume's terminal states are completed, failed and canceled.
Tradeoffs
A fake can drift from the API. Keep one smoke test against the real API with a cheap job, and treat the fake as a guard for your loop logic, not for Sume's behavior.
Run the test in CI with no secrets at all. Since the base URL is a parameter, the same function that talks to the fake server talks to https://api.sume.com/v1 in production, with the Authorization: Bearer header added. That is the whole seam, and it keeps the test honest about real HTTP behavior.
Sources
Related posts
More in Developers
- Fall back to a second image model after three 502s: a Python breaker
Sume returns 502 when an image job fails inside the wait budget. Count them, switch to a second model id after three, and use a new idempotency key per model.
- Four 16:9 thumbnail options in one call: n=4 cost, and why n=5 fails
Set n to 4 for four 1536x864 options in one GPT Image 2.5 call: about $0.0420 at medium on Sume. n=5 returns a 400, because every model caps at 4.
- Gemini CLI and the hosted Sume server: do not rely on env in headers
Add the hosted Sume server to Gemini CLI with httpUrl and a bearer header. Gemini expands env vars only in the env block; set a timeout above jobs_wait.
- Gemini Omni 4K on Sume: send 4K or 4k, 10 s max, 16:9 or 9:16
Sume's gemini-omni-flash-1.1 takes 4K, and lowercase 4k works as an alias. Requests run 3 to 10 seconds, 16:9 or 9:16. Python body builder.
Written by Sume