A Sume poll returned HTML: guard the JSON parse in Python

A proxy or edge can answer a Sume poll with an HTML page. A tested stdlib Python helper that returns None for non-JSON bodies so your loop retries, not crashes.

3 min readSume
All posts

If a poll to the Sume API passes through a proxy or CDN that times out, the answer can be an HTML error page, and calling json() on it crashes your loop. Wrap the parse so a non-JSON body becomes a retryable result, and keep the job id so the next read goes to the same job.

The helper below uses only the standard library. Against a local server that returned a 524 with an HTML body, it returned (524, None), which is the signal to retry.

The helper

It returns a status code and a dict, or a status code and None when the body is not JSON. Transport errors return (None, None).

Two design choices are worth noting. The helper catches HTTPError separately, because urllib raises it for every non-2xx response and the body of that response is still readable, which is where Sume's JSON error envelope lives. And it treats a 200 whose body is not JSON as a transport failure, since a proxy can rewrite a response without changing its status. In both cases the caller sees None and the original status, so the retry rule is a single if statement and the status code is available for your logs.

import json, urllib.error, urllib.request

def read_json(url: str, key: str):
    """Return (status, dict) or (status, None) when the body is not Sume JSON."""
    req = urllib.request.Request(url, headers={"Authorization": f"Bearer {key}"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            return resp.status, json.loads(resp.read())
    except urllib.error.HTTPError as err:        # non-2xx still has a body
        raw = err.read()
        try:
            return err.code, json.loads(raw)
        except ValueError:                       # HTML page from an edge
            return err.code, None
    except (urllib.error.URLError, TimeoutError, ValueError):
        return None, None                        # transport failure or non-JSON 200

if __name__ == "__main__":
    import threading, http.server
    class H(http.server.BaseHTTPRequestHandler):
        def do_GET(self):
            self.send_response(524); self.send_header("Content-Type", "text/html")
            self.end_headers(); self.wfile.write(b"<html>timeout</html>")
        def log_message(self, *a): pass
    srv = http.server.HTTPServer(("127.0.0.1", 0), H)
    threading.Thread(target=srv.serve_forever, daemon=True).start()
    print(read_json(f"http://127.0.0.1:{srv.server_port}/v1/jobs/job_1", "k"))

How the loop should use it

Treat None as retry, not as a failure of the job. A job whose submit succeeded exists and is billed regardless of what a poll returns.

Helper results and what the poll loop does (read 2026-10-06, from the tested sample)
Return valueLoop action
(200, dict)Read terminal and next_action
(409, dict)Read the error code in the dict
(524, None)Retry with backoff
(None, None)Retry; transport failure or non-JSON 200

Why this matters for cost

The tempting reaction to a confusing response is to submit again. Do not. Sume's docs say that after a 2xx submit you poll, and a repeated submit with the same Idempotency-Key returns the original job rather than starting another.

Tradeoffs

Swallowing every parse error can hide a real bug, such as a wrong base URL that serves HTML for everything. Log the status and the first bytes of the body on retry, and fail after a fixed number of consecutive non-JSON answers.

Add a counter of consecutive None results and give up, with the job id logged, after a limit you choose. That separates a brief edge hiccup from a long outage, and it leaves the job itself alone: you can resume polling the same id later, since nothing here cancels it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume