What to measure for Sume API jobs: metrics, labels and alerts

A metrics plan for code that calls the Sume API: submit outcome, queue wait, time to terminal, error code and webhook gap, with low-cardinality labels.

5 min readSume
All posts

Five measurements cover most of what goes wrong in a Sume integration: submit outcome, time spent queued, time to terminal, terminal error code, and the gap between a job finishing and your webhook handler seeing it. All of them come from fields the API already returns.

The measurements

Use the job status vocabulary as labels, because it is a closed set: queued, processing, completed, failed, canceled. Use error.code for submit failures. Keep job ids and request ids out of metric labels and in logs, where high cardinality is fine.

Sume integration metrics (read 2026-10-03)
MetricSource fieldWhy it matters
Submit result countHTTP status and error.codeSeparates 402, 429 and 503 causes
Queue waitjob.queued to job.started eventsShows plan concurrency pressure
Time to terminaljob.created to job.completed or job.failedSets your client deadline
Terminal outcome countcompleted, failed, canceledFailure rate by model
Webhook gapterminal time minus handler receiptSpots missed deliveries

Alerts that are worth waking up for

Alert on a rising share of 402 insufficient_credits, since retrying never helps and it means spend or top-up drifted. Alert on any sustained queue_full, which means you submit faster than the plan processes. Alert on webhook gaps, then fall back to polling status_url; the docs describe delivery as an optimization, never the only recovery path.

Where the timestamps live

GET /v1/jobs/:id/events returns a public timeline with job.created, job.queued, job.started, generation.submitted, and the terminal events, plus webhook.delivery. Read it for one slow job rather than polling it for every job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume