Prometheus counters for Sume API errors, split by code and status

Wrap every Sume call in a counter and histogram labeled by route template, status and error.code, never by job id or request id, then alert on retryable rates.

5 min readSume
All posts

To watch a Sume integration in Prometheus, wrap the HTTP call once and record a counter labeled route, status and code, plus a latency histogram labeled route. Take code from error.code in the response body, because the HTTP status alone cannot tell 429 rate_limited from 429 queue_full, and the right reaction to each is different. Never put a job id, run id or request_id in a label.

The error envelope is documented in Formats: errors and spend: a lowercase code token you can branch on, a request_id that is also sent as the x-sume-request-id header, and on that surface retryable, retry_after_seconds and next_action. Treat the code set as open, since new codes can appear.

What does the wrapper look like?

The route label is a template you choose (jobs_status), not the URL, so a million job ids still produce one series.

import os, time, requests
from prometheus_client import Counter, Histogram, start_http_server

CALLS = Counter("sume_api_calls_total", "Sume API calls", ["route", "status", "code"])
SECONDS = Histogram("sume_api_seconds", "Sume API latency", ["route"])

def sume(method, route, path, **kw):
    t0 = time.monotonic()
    r = requests.request(method, "https://api.sume.com" + path, timeout=30,
        headers={"x-api-key": os.environ["SUME_API_KEY"]}, **kw)
    SECONDS.labels(route).observe(time.monotonic() - t0)
    code = "ok"
    if r.status_code >= 400:
        try:
            code = r.json()["error"]["code"]
        except (ValueError, KeyError, TypeError):
            code = "unparsed"
    CALLS.labels(route, str(r.status_code), code).inc()
    return r

if __name__ == "__main__":
    start_http_server(9108)
    r = sume("GET", "jobs_status", "/v1/jobs/job_123/status")
    print(r.status_code, r.headers.get("x-sume-request-id"))

Which alerts are worth writing?

Page on rates, not on single errors. The table pairs a code with its documented meaning and a sensible alert shape.

Alert candidates by Sume error code, from Sume docs read 2026-10-04
code (status)What it meansAlert on
rate_limited (429)The write or read budget for the key is spentSustained rate above 1 percent of calls; check route mix
queue_full (429)Concurrency plus queue capacity is fullAny; you are submitting faster than jobs finish
insufficient_credits (402)The wallet cannot fund the jobAny; page a human, retrying cannot help
insufficient_scope (403)The key lacks a scopeAny; a retry loop here is expensive
studio_agent_upstream_unavailable (503)A Sume-side outage on a Format callRate over 5 minutes; retries are safe with the same key
unparsedBody was not the envelopeAny; usually a proxy in front of you

Why keep ids out of labels?

Every distinct label value creates a new time series, and a request_id is unique per call. Put ids in logs and traces, where you can search them, and keep the metric labels to values with a small known set. Quote the x-sume-request-id from a failing response to Sume support; do not include API keys, signing secrets, or raw media URLs when you do.

What about polling cost?

The same counter shows how much of your request budget polling uses. Reads and writes are budgeted separately per key by plan, and every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, so a gauge of ratelimit-remaining per route tells you how close you are before a 429 arrives. See Errors and rate limits for the header list.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume