Hourly synthetic monitor for an image API on a cheap Sume model
A scheduled script that makes one real image call, checks status, URL and latency, and exits non-zero on failure. Budget it from the endpoint's cost_usd.

A synthetic monitor for an image pipeline is a scheduled script that makes one real POST /v1/images call with a small prompt, checks that the answer is a 200 with an image URL inside a latency budget, and exits non-zero otherwise so your scheduler pages you. Run it hourly on the cheapest model that supports the parameters you use in production, and cap the monthly bill from the model's cost_usd in GET /v1/images/models/{id}/endpoints.
A real call matters because a health check cannot tell you that generation works end to end. It does cost money each time, which is why the model and schedule are chosen to keep the spend small and known.
What does it cost to run?
Read the per-image price from the endpoint, then multiply. Sume documents cost_usd as the charge per unit with margin applied.
| Input | Where it comes from | Example use |
|---|---|---|
| Price per image | pricing[].cost_usd on the endpoint | Read at deploy time |
| Runs per day | Your schedule | 24 for hourly |
| Days per month | Calendar | 30 for a rough figure |
| Monthly spend | Price x runs x days | Compute it, then alert if above budget |
What does the script check?
It treats a 202 as a failure for this purpose: a tiny image should finish inside the 30 second sync wait, so a 202 is a latency signal worth paging on.
import os, sys, time, requests
MODEL = os.environ.get("MONITOR_MODEL", "qwen/qwen-image")
BUDGET_SECONDS = 25
def main():
t0 = time.time()
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
json={"model": MODEL, "prompt": "A plain blue square", "n": 1},
timeout=40,
)
took = time.time() - t0
if r.status_code != 200:
print(f"FAIL status {r.status_code}: {r.text[:200]}")
return 1
url = r.json()["data"][0]["url"]
if not url.startswith("https://") or took > BUDGET_SECONDS:
print(f"FAIL url={url} took={took:.1f}s")
return 1
print(f"OK {took:.1f}s cost={r.json()['usage']['cost']}")
return 0
if __name__ == "__main__":
sys.exit(main())How do I schedule it?
Any scheduler that alerts on a non-zero exit works: cron with a mail-on-failure wrapper, a scheduled CI workflow or your platform's job runner. Store the key in that system's secret store and give the monitor its own key so you can see its usage separately.
What should happen when it fails?
Read the error code and next_action before deciding who gets paged. A fix_input failure is on your side; a retryable one may clear on the next run. The 502 error post walks through the envelope, and retryable status codes covers the rest. Failed generations are not billed, so a failing monitor does not run up a bill.
Sources
Related posts
More in Developers
- End-user id on jobs: OpenAI safety identifier vs Sume metadata
OpenAI's Realtime guide asks for an OpenAI-Safety-Identifier header. Sume stores caller metadata on the job but does not send it to the provider. Use both.
- Temporal Paygo starter: submit and poll a Sume job
Temporal's Paygo plan has a $0 monthly minimum. A first workflow can submit a Sume job in one Activity and poll its status in a second. Python code included.
- Test a faster-and-cheaper claim with your own timings and usage.cost
Luma's news page says Ray3.14 is 4x faster and 3x cheaper. A short Python script turns your own job timings and usage.cost values into two ratios you can trust.
- Test MAI-Transcribe-2-Streaming's accuracy claim on your own audio
Microsoft says its streaming transcriber ranks first on Artificial Analysis. A 10-clip test on your own audio, with a word error rate script and Sume STT.
Written by Sume