AI image prompt regression test: check size, ratio and alpha in CI

Before you swap or upgrade an image model, run a fixed prompt set through Sume and assert pixel size, aspect ratio and alpha with Pillow. Python test included.

5 min readSume
All posts

To regression-test an image prompt before switching models, keep a small file of prompts with the properties you must not lose, call the Sume image endpoint for each, and assert on the downloaded file with Pillow: pixel dimensions, aspect ratio, alpha channel and file format. You cannot assert on the picture being good, but you can catch the failures that break a layout: a wrong ratio, a missing transparent background, a format your CMS rejects.

This matters now because image model ids change often. ChatGPT Image 2 stays selectable next to ChatGPT Image 2.5 on Sume, and new catalog rows appear regularly, so a swap should be a tested change rather than a guess.

What can a test assert without a human?

Everything that is a property of the file. The checks below are cheap and deterministic once the image is downloaded.

Checks per prompt, read 2026-10-04
CheckHowCatches
Aspect ratiowidth / height within 2% of the requestA model that snaps to its nearest native ratio
Minimum pixelsmin(width, height) >= floorA tier silently lower than you priced for
Alphaimg.mode == "RGBA" and some alpha below 255A background that came back opaque
Formatimg.format equals the one you requestedA provider-selected format your pipeline rejects
StatusHTTP 200, not 202Slow configurations that now exceed the 30 s wait

What does the harness look like?

The test reads cases.json, a list of objects with model, prompt, aspect_ratio, and optional background. Remember that background is only accepted by ChatGPT Image 2.5; other models return 400 unsupported_parameter, which is itself a useful assertion when you change the model in a case.

import io, json, os, requests
from PIL import Image

HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(case):
    r = requests.post("https://api.sume.com/v1/images", headers=HEAD,
                      json=case, timeout=90)
    assert r.status_code == 200, (case["model"], r.status_code, r.text[:200])
    raw = requests.get(r.json()["data"][0]["url"], timeout=60).content
    return Image.open(io.BytesIO(raw))

def test_cases():
    for case in json.load(open("cases.json")):
        img = run(case)
        w, h = img.size
        a, b = map(int, case["aspect_ratio"].split(":"))
        assert abs(w / h - a / b) < 0.02 * a / b, (case, img.size)
        if case.get("background") == "transparent":
            assert img.mode == "RGBA" and img.getchannel("A").getextrema()[0] < 255

if __name__ == "__main__":
    test_cases()
    print("ok")

How should I keep the cost of the suite low?

Run each case at low quality where the model accepts it, with n left at 1, and keep the suite to a dozen prompts. Sume bills a completed generation in full and does not bill failed or cancelled ones, so an assertion failure on your side still costs the image you downloaded. Check each model's list price in GET /v1/images/models/{id}/endpoints before you add it to the suite.

When should the suite run?

Run it when you change a model id, when the catalog adds a row you intend to adopt, and weekly on a schedule to catch drift in what a given id returns. Treat a ratio or alpha failure as a blocker and a content difference as a human review item, since the test cannot judge it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume