Hourly synthetic monitor for an image API on a cheap Sume model

A scheduled script that makes one real image call, checks status, URL and latency, and exits non-zero on failure. Budget it from the endpoint's cost_usd.

5 min readSume
All posts

A synthetic monitor for an image pipeline is a scheduled script that makes one real POST /v1/images call with a small prompt, checks that the answer is a 200 with an image URL inside a latency budget, and exits non-zero otherwise so your scheduler pages you. Run it hourly on the cheapest model that supports the parameters you use in production, and cap the monthly bill from the model's cost_usd in GET /v1/images/models/{id}/endpoints.

A real call matters because a health check cannot tell you that generation works end to end. It does cost money each time, which is why the model and schedule are chosen to keep the spend small and known.

What does it cost to run?

Read the per-image price from the endpoint, then multiply. Sume documents cost_usd as the charge per unit with margin applied.

Monitor budget formula, read 2026-10-04
InputWhere it comes fromExample use
Price per imagepricing[].cost_usd on the endpointRead at deploy time
Runs per dayYour schedule24 for hourly
Days per monthCalendar30 for a rough figure
Monthly spendPrice x runs x daysCompute it, then alert if above budget

What does the script check?

It treats a 202 as a failure for this purpose: a tiny image should finish inside the 30 second sync wait, so a 202 is a latency signal worth paging on.

import os, sys, time, requests

MODEL = os.environ.get("MONITOR_MODEL", "qwen/qwen-image")
BUDGET_SECONDS = 25

def main():
    t0 = time.time()
    r = requests.post(
        "https://api.sume.com/v1/images",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
        json={"model": MODEL, "prompt": "A plain blue square", "n": 1},
        timeout=40,
    )
    took = time.time() - t0
    if r.status_code != 200:
        print(f"FAIL status {r.status_code}: {r.text[:200]}")
        return 1
    url = r.json()["data"][0]["url"]
    if not url.startswith("https://") or took > BUDGET_SECONDS:
        print(f"FAIL url={url} took={took:.1f}s")
        return 1
    print(f"OK {took:.1f}s cost={r.json()['usage']['cost']}")
    return 0

if __name__ == "__main__":
    sys.exit(main())

How do I schedule it?

Any scheduler that alerts on a non-zero exit works: cron with a mail-on-failure wrapper, a scheduled CI workflow or your platform's job runner. Store the key in that system's secret store and give the monitor its own key so you can see its usage separately.

What should happen when it fails?

Read the error code and next_action before deciding who gets paged. A fix_input failure is on your side; a retryable one may clear on the next run. The 502 error post walks through the envelope, and retryable status codes covers the rest. Failed generations are not billed, so a failing monitor does not run up a bill.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume