Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.

To contract-test your Sume integration, save a real response from your own account as a fixture, load Sume's OpenAPI document, and validate the fixture against the response schema with jsonschema. The one non-obvious step is nullable fields: the spec is OpenAPI 3.0.3, which marks them with nullable: true, and a plain JSON Schema validator rejects every null until you convert them to a type union.
I ran the sample below with jsonschema 4.26 against the repository's OpenAPI snapshot. A fixture passed, and a copy with one required field deleted and one enum value changed produced exactly two errors. Without the conversion, the same valid fixture failed on every nullable field.
Why test responses against the spec at all?
Your parsing code is written against a response shape you saw once. A contract test turns that memory into a check that runs in CI: when your fixture stops matching the published schema, or when you refresh the schema and an old fixture no longer fits, you find out before a customer does. It is also a cheap way to catch your own mistakes, such as reading status when you meant sume_status.
Sume's docs say the production API serves its live schema at https://api.sume.com/reference/json, with a snapshot in the local docs preview, so a nightly job can download the live document and re-run the tests against stored fixtures.
| Question | Covered by a schema test? | Better tool |
|---|---|---|
| Did a required field disappear from the response? | Yes | This test |
| Did an enum gain or lose a value your code switches on? | Yes, for values in your fixtures | This test, plus a default branch in code |
| Is a nullable field null where your code assumes a value? | Partly | Type checks and code review |
| Does the job finish, and with the right pixels? | No | A real end-to-end run |
| Is my webhook signature check correct? | No | Unit tests on the verifier |
What does the test look like?
to_2020 walks the document and rewrites each nullable: true into a type union, adding None to an enum when the field has one. validator builds a validator for one path, method and status by pointing at that response's schema and attaching the converted components, so $ref entries such as #/components/schemas/JobStatusResponse resolve. The fixture is a status response you recorded yourself.
import json, pathlib
from jsonschema import Draft202012Validator
SPEC = json.loads(pathlib.Path("openapi.json").read_text()) # from api.sume.com/reference/json
def to_2020(node): # OpenAPI 3.0 "nullable" becomes a JSON Schema type union
if isinstance(node, list):
return [to_2020(v) for v in node]
if not isinstance(node, dict):
return node
out = {k: to_2020(v) for k, v in node.items() if k != "nullable"}
if node.get("nullable") and "type" in out:
out["type"] = [out["type"], "null"]
if "enum" in out:
out["enum"] = [*out["enum"], None]
return out
def validator(path, method, status="200"):
schema = SPEC["paths"][path][method]["responses"][status]["content"]["application/json"]["schema"]
return Draft202012Validator({**to_2020(schema), "components": to_2020(SPEC["components"])})
def load():
return json.loads(pathlib.Path("fixtures/status.json").read_text())
def test_recorded_status_matches_the_spec():
errors = [e.message for e in validator("/v1/jobs/{id}/status", "get").iter_errors(load())]
assert errors == []
def test_drift_is_caught():
body = load()
del body["data"]["terminal"]
body["data"]["sume_status"] = "done"
assert len(list(validator("/v1/jobs/{id}/status", "get").iter_errors(body))) == 2How do you get good fixtures?
Record them from your own account instead of writing them by hand: a queued job, a processing job, a completed one and a failed one, each saved as JSON with ids and URLs left as they came. The status schema marks fields like cancel_url, next_poll_after_seconds and queue_position as nullable, and the field descriptions say next_poll_after_seconds is null for terminal jobs, so a queued fixture and a completed fixture exercise different branches of the schema. A fixture that is all nulls, or all values, can hide a mistake that the other state would reveal.
Save error bodies too. The OpenAPI lists 400, 401, 404, 429, 500 and 503 responses per operation, and validator(path, method, "429") checks a recorded error body the same way, provided that response declares a JSON body. When I fed a bare error with only code, message and request_id to the 429 validator, it reported six missing required fields: category, stage, retryable, retry_after_seconds, public_reason and next_action. That is the kind of shortcut a hand-written fixture would hide. A test that your error parser reads error.code from a recorded 429 fixture is worth more than any amount of reading.
- Commit fixtures, and commit the OpenAPI snapshot they were checked against.
- Refresh the snapshot on purpose, in its own change, so a schema diff is easy to review.
- Keep fixtures free of keys, signed URLs and real prompts.
- Assert on required fields you actually read, not on the whole shape.
- Add a default branch for unknown enum values in the code under test.
Where does this approach stop?
A schema check proves shape, not behaviour, and it only knows about the responses you have recorded. A new enum value in the live spec will not fail a stored fixture, so also diff the downloaded schema against your snapshot and read the changes. The spec's own limits apply too: it is OpenAPI 3.0.3, so tools that expect 3.1 or newer need the same kind of conversion.
Used this way the test is small, fast and offline, and it covers the quiet failure that hurts: a field you depend on changing shape without anything crashing.
Sources
Related posts
More in Developers
- DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.
- Dub one Short into 8 languages: Python fan-out and the total cost
Detach and transcribe once, then run one TTS job and one render per language. A Python fan-out and the per-Short bill, from Sume's catalog rates.
- FLUX 3 bounding box to a mask_url: Python region edit on Sume
FLUX 3 Image boxes use [top, left, bottom, right] on a 0-1000 grid. Convert one to an RGBA mask with Pillow and run the region edit on Sume's GPT Image 2.5.
- FLUX 3 Image on OpenRouter: n=1, seed, base64 vs Sume
OpenRouter lists FLUX.3 Image with one image per call, a seed and base64 PNG output. How each differs from Sume's POST /v1/images, where FLUX 3 is not listed.
Written by Sume