Datadog DogStatsD metrics for Sume jobs: tags that stay small

Count Sume submits, terminal statuses and durations with DogStatsD, tagged by route, status and code, and keep job ids and request ids out of the tag set.

5 min readSume
All posts

To see Sume jobs in Datadog, send three metrics through DogStatsD: a counter for submits, a counter for terminal outcomes and a timing for submit-to-terminal duration, tagged with a small fixed set such as route, status and code. Never tag with a job id, run id or request id. The datadogpy README shows the pattern: initialize the client, then use statsd against the local Agent on UDP 127.0.0.1:8125 (read 2026-10-04).

Keep the id in logs, where you can search it, and the tags to values with a known, small set. Each new tag value creates a new series for a custom metric, so an id tag turns one metric into one series per job.

What does the instrumentation look like?

Wrap the one place you read a terminal status. Statuses come from Jobs and results: completed, failed or canceled.

import time
from datadog import initialize, statsd

initialize(statsd_host="127.0.0.1", statsd_port=8125)

def on_submit(route: str) -> float:
    statsd.increment("sume.job.submitted", tags=[f"route:{route}"])
    return time.monotonic()

def on_terminal(route: str, status: str, started: float, code: str = "none") -> None:
    tags = [f"route:{route}", f"status:{status}", f"code:{code}"]
    statsd.increment("sume.job.terminal", tags=tags)
    statsd.histogram("sume.job.seconds", time.monotonic() - started, tags=tags[:2])

def on_api_error(route: str, http_status: int, code: str) -> None:
    statsd.increment("sume.api.error",
                     tags=[f"route:{route}", f"http:{http_status}", f"code:{code}"])

if __name__ == "__main__":
    t0 = on_submit("image_generate")
    on_terminal("image_generate", "completed", t0)
    on_api_error("jobs_status", 429, "rate_limited")

Which tags are worth having?

Tag choices for Sume metrics, from Sume docs read 2026-10-04
TagValuesKeep?
routeA template you name, such as image_generateYes; a handful of values
statuscompleted, failed, canceledYes; three values
codeerror.code tokens such as rate_limitedYes; the set is small but open, so alert on unknown values
httpStatus codes you seeYes
job_id, run_id, request_idUnique per callNo; put them in logs
prompt textUnboundedNo

What should the monitors say?

Alert on rates over a window, not on single events. A sustained rate_limited ratio means your poll interval is too short; read the retry-after header and the next_poll_after_seconds hint in the status body. insufficient_credits on any call is a funding problem for a person, and a retry loop cannot fix it.

Where does this fall short?

  • UDP is fire and forget, so a metric can be dropped without an error in your code.
  • A failed job is not an API error: the status read returned 200. Count it from the terminal status, not from the HTTP code.
  • Use a webhook-driven counter as well if jobs can finish while no process is polling.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume