What to measure for Sume API jobs: metrics, labels and alerts
A metrics plan for code that calls the Sume API: submit outcome, queue wait, time to terminal, error code and webhook gap, with low-cardinality labels.

Five measurements cover most of what goes wrong in a Sume integration: submit outcome, time spent queued, time to terminal, terminal error code, and the gap between a job finishing and your webhook handler seeing it. All of them come from fields the API already returns.
The measurements
Use the job status vocabulary as labels, because it is a closed set: queued, processing, completed, failed, canceled. Use error.code for submit failures. Keep job ids and request ids out of metric labels and in logs, where high cardinality is fine.
| Metric | Source field | Why it matters |
|---|---|---|
| Submit result count | HTTP status and error.code | Separates 402, 429 and 503 causes |
| Queue wait | job.queued to job.started events | Shows plan concurrency pressure |
| Time to terminal | job.created to job.completed or job.failed | Sets your client deadline |
| Terminal outcome count | completed, failed, canceled | Failure rate by model |
| Webhook gap | terminal time minus handler receipt | Spots missed deliveries |
Alerts that are worth waking up for
Alert on a rising share of 402 insufficient_credits, since retrying never helps and it means spend or top-up drifted. Alert on any sustained queue_full, which means you submit faster than the plan processes. Alert on webhook gaps, then fall back to polling status_url; the docs describe delivery as an optimization, never the only recovery path.
Where the timestamps live
GET /v1/jobs/:id/events returns a public timeline with job.created, job.queued, job.started, generation.submitted, and the terminal events, plus webhook.delivery. Read it for one slow job rather than polling it for every job.
Sources
Related posts
More in Developers
- Which Sume API errors should page an engineer: route by category
Route Sume API failures by category: fix-the-input errors go to the caller, quota to finance, queue to a retry, and only internal or unexpected 5xx to on-call.
- Crash-safe Sume submit: write the intent row and key first
If your process dies after a Sume submit but before storing the job id, a pre-written intent row and Idempotency-Key let the retry return the original job.
- Zod 4 discriminated union for Sume job and run webhooks (TypeScript)
Parse Sume job.* and format.run.terminal webhooks with one Zod 4 discriminatedUnion: typed branches, degraded runs, oversized receipts. Tested with Zod 4.
- Zod 4 toJSONSchema to Sume output_schema: nullable, not optional
z.toJSONSchema works for a Sume Format output_schema if you use nullable instead of optional. A tested table of what passes and what the validator rejects.
Written by Sume