Prometheus counters for Sume API errors, split by code and status
Wrap every Sume call in a counter and histogram labeled by route template, status and error.code, never by job id or request id, then alert on retryable rates.

To watch a Sume integration in Prometheus, wrap the HTTP call once and record a counter labeled route, status and code, plus a latency histogram labeled route. Take code from error.code in the response body, because the HTTP status alone cannot tell 429 rate_limited from 429 queue_full, and the right reaction to each is different. Never put a job id, run id or request_id in a label.
The error envelope is documented in Formats: errors and spend: a lowercase code token you can branch on, a request_id that is also sent as the x-sume-request-id header, and on that surface retryable, retry_after_seconds and next_action. Treat the code set as open, since new codes can appear.
What does the wrapper look like?
The route label is a template you choose (jobs_status), not the URL, so a million job ids still produce one series.
import os, time, requests
from prometheus_client import Counter, Histogram, start_http_server
CALLS = Counter("sume_api_calls_total", "Sume API calls", ["route", "status", "code"])
SECONDS = Histogram("sume_api_seconds", "Sume API latency", ["route"])
def sume(method, route, path, **kw):
t0 = time.monotonic()
r = requests.request(method, "https://api.sume.com" + path, timeout=30,
headers={"x-api-key": os.environ["SUME_API_KEY"]}, **kw)
SECONDS.labels(route).observe(time.monotonic() - t0)
code = "ok"
if r.status_code >= 400:
try:
code = r.json()["error"]["code"]
except (ValueError, KeyError, TypeError):
code = "unparsed"
CALLS.labels(route, str(r.status_code), code).inc()
return r
if __name__ == "__main__":
start_http_server(9108)
r = sume("GET", "jobs_status", "/v1/jobs/job_123/status")
print(r.status_code, r.headers.get("x-sume-request-id"))Which alerts are worth writing?
Page on rates, not on single errors. The table pairs a code with its documented meaning and a sensible alert shape.
| code (status) | What it means | Alert on |
|---|---|---|
| rate_limited (429) | The write or read budget for the key is spent | Sustained rate above 1 percent of calls; check route mix |
| queue_full (429) | Concurrency plus queue capacity is full | Any; you are submitting faster than jobs finish |
| insufficient_credits (402) | The wallet cannot fund the job | Any; page a human, retrying cannot help |
| insufficient_scope (403) | The key lacks a scope | Any; a retry loop here is expensive |
| studio_agent_upstream_unavailable (503) | A Sume-side outage on a Format call | Rate over 5 minutes; retries are safe with the same key |
| unparsed | Body was not the envelope | Any; usually a proxy in front of you |
Why keep ids out of labels?
Every distinct label value creates a new time series, and a request_id is unique per call. Put ids in logs and traces, where you can search them, and keep the metric labels to values with a small known set. Quote the x-sume-request-id from a failing response to Sume support; do not include API keys, signing secrets, or raw media URLs when you do.
What about polling cost?
The same counter shows how much of your request budget polling uses. Reads and writes are budgeted separately per key by plan, and every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, so a gauge of ratelimit-remaining per route tells you how close you are before a 429 arrives. See Errors and rate limits for the header list.
Sources
Related posts
More in Developers
- Export Sume job counts by status to Prometheus with a Python gauge
A small exporter pages GET /v1/jobs with next_cursor and sets a prometheus_client Gauge labelled by status, so a dashboard shows queued and failed jobs.
- Push or poll for a finished render: listen, webhook or jobs_wait
MCP 2026-07-28 adds subscriptions/listen. For a render that takes minutes, compare a listen stream, a signed webhook and jobs_wait, with a Python verifier.
- Pydantic AI slot leak vs Sume queue_full: tell them apart
Pydantic AI v2.53.0 fixed a streamed-request concurrency slot leak. A client limiter is not Sume's workspace queue_full 429, and each needs its own handling.
- Pydantic AI Workspace sandboxes: fetch Sume artifacts
Pydantic AI's Workspace abstraction runs tools locally or in sandboxes. Inside a sandbox, fetch Sume artifact URLs, and give Sume inputs as public HTTPS URLs.
Written by Sume