Export Sume job counts by status to Prometheus with a Python gauge
A small exporter pages GET /v1/jobs with next_cursor and sets a prometheus_client Gauge labelled by status, so a dashboard shows queued and failed jobs.

Page GET /v1/jobs?limit=100 until data.next_cursor is absent, count job.status in each page, and set Gauge('sume_jobs', ..., labelnames=['status']).labels(status=s).set(n). Expose the gauge with start_http_server. The counts describe the window of jobs you scan, not your lifetime total.
The list contract
The job list takes limit (1 to 100), scope, status, type and starting_after. The response puts jobs in data.jobs and returns next_cursor only when more remain; the schema calls its absence the loop terminator.
| Piece | Value |
|---|---|
| Page size | limit 1 to 100 |
| Status values | queued, processing, completed, failed, canceled |
| Next page | pass next_cursor back as starting_after |
| Last page | next_cursor is absent |
Exporter
The prometheus_client docs create a gauge with labelnames, set a labelled value with .labels(...).set(...), and serve metrics with start_http_server. The loop sets every known status each cycle, so a status that drops to zero reports 0 instead of keeping its last value.
import asyncio, json, os, urllib.parse, urllib.request
from prometheus_client import Gauge, start_http_server
STATUSES = ["queued", "processing", "completed", "failed", "canceled"]
JOBS = Gauge("sume_jobs", "Sume jobs in the scanned window, by status", labelnames=["status"])
MAX_PAGES = 5
def fetch(cursor):
q = {"limit": "100"}
if cursor: q["starting_after"] = cursor
req = urllib.request.Request("https://api.sume.com/v1/jobs?" + urllib.parse.urlencode(q),
headers={"x-api-key": os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]
def scan():
counts, cursor = dict.fromkeys(STATUSES, 0), None
for _ in range(MAX_PAGES):
page = fetch(cursor)
for job in page["jobs"]:
counts[job["status"]] = counts.get(job["status"], 0) + 1
cursor = page.get("next_cursor")
if not cursor: break
return counts
async def main():
start_http_server(8000)
while True:
for status, n in (await asyncio.to_thread(scan)).items():
JOBS.labels(status=status).set(n)
await asyncio.sleep(60)
asyncio.run(main())Cost and cadence
Each scrape cycle costs up to five read requests. Reads have their own, much larger budget than writes (40 times the write budget per the rate-limit docs), but keep the interval at a minute or more and the page cap low.
What to alert on
Alert on sume_jobs{status="queued"} staying high while processing is at your plan's concurrency, which means you are waiting on a slot rather than on the model. Per the admission docs, concurrency is 1 on Free, 4 on Pro, 8 on Startup and 20 on Scale.
Sources
Related posts
More in Developers
- Push or poll for a finished render: listen, webhook or jobs_wait
MCP 2026-07-28 adds subscriptions/listen. For a render that takes minutes, compare a listen stream, a signed webhook and jobs_wait, with a Python verifier.
- Pydantic AI slot leak vs Sume queue_full: tell them apart
Pydantic AI v2.53.0 fixed a streamed-request concurrency slot leak. A client limiter is not Sume's workspace queue_full 429, and each needs its own handling.
- Pydantic AI Workspace sandboxes: fetch Sume artifacts
Pydantic AI's Workspace abstraction runs tools locally or in sandboxes. Inside a sandbox, fetch Sume artifact URLs, and give Sume inputs as public HTTPS URLs.
- pytest and respx: fake a Sume job that is queued, then completed
Test your Sume poll loop with respx: a queued, a processing and a completed response in order, with sleep injected so the test finishes in milliseconds.
Written by Sume