Datadog DogStatsD metrics for Sume jobs: tags that stay small
Count Sume submits, terminal statuses and durations with DogStatsD, tagged by route, status and code, and keep job ids and request ids out of the tag set.

To see Sume jobs in Datadog, send three metrics through DogStatsD: a counter for submits, a counter for terminal outcomes and a timing for submit-to-terminal duration, tagged with a small fixed set such as route, status and code. Never tag with a job id, run id or request id. The datadogpy README shows the pattern: initialize the client, then use statsd against the local Agent on UDP 127.0.0.1:8125 (read 2026-10-04).
Keep the id in logs, where you can search it, and the tags to values with a known, small set. Each new tag value creates a new series for a custom metric, so an id tag turns one metric into one series per job.
What does the instrumentation look like?
Wrap the one place you read a terminal status. Statuses come from Jobs and results: completed, failed or canceled.
import time
from datadog import initialize, statsd
initialize(statsd_host="127.0.0.1", statsd_port=8125)
def on_submit(route: str) -> float:
statsd.increment("sume.job.submitted", tags=[f"route:{route}"])
return time.monotonic()
def on_terminal(route: str, status: str, started: float, code: str = "none") -> None:
tags = [f"route:{route}", f"status:{status}", f"code:{code}"]
statsd.increment("sume.job.terminal", tags=tags)
statsd.histogram("sume.job.seconds", time.monotonic() - started, tags=tags[:2])
def on_api_error(route: str, http_status: int, code: str) -> None:
statsd.increment("sume.api.error",
tags=[f"route:{route}", f"http:{http_status}", f"code:{code}"])
if __name__ == "__main__":
t0 = on_submit("image_generate")
on_terminal("image_generate", "completed", t0)
on_api_error("jobs_status", 429, "rate_limited")Which tags are worth having?
| Tag | Values | Keep? |
|---|---|---|
| route | A template you name, such as image_generate | Yes; a handful of values |
| status | completed, failed, canceled | Yes; three values |
| code | error.code tokens such as rate_limited | Yes; the set is small but open, so alert on unknown values |
| http | Status codes you see | Yes |
| job_id, run_id, request_id | Unique per call | No; put them in logs |
| prompt text | Unbounded | No |
What should the monitors say?
Alert on rates over a window, not on single events. A sustained rate_limited ratio means your poll interval is too short; read the retry-after header and the next_poll_after_seconds hint in the status body. insufficient_credits on any call is a funding problem for a person, and a retry loop cannot fix it.
Where does this fall short?
- UDP is fire and forget, so a metric can be dropped without an error in your code.
- A failed job is not an API error: the status read returned 200. Count it from the terminal status, not from the HTTP code.
- Use a webhook-driven counter as well if jobs can finish while no process is polling.
Sources
Related posts
More in Developers
- Demand Gen copy limits: 40-character headlines, one at 30 or fewer
Demand Gen allows 40-character headlines (one must be 30 or fewer), 90-character descriptions and 10-60 second videos. A checker script plus the Sume lengths.
- Design a render tool for a stateless MCP server: job ids as arguments
MCP 2026-07-28 removes protocol sessions. A render tool stays correct if its state lives in a job id the client passes back, as Sume jobs do.
- Cline oversized MCP result cache: keep Sume job results small
Cline's SDK now caches MCP results over 16 MiB and gives a preview plus a cache URI. Keep Sume job output small with jobs_wait include_results.
- Cut a 30-second hook from a video's audio: detach with a range
Pull just seconds 45 to 75 of a Sume-hosted video as a mono mp3 with POST /v1/audio-detach and range. Source and output caps, price and refusals.
Written by Sume