Group failed Sume jobs by error category with curl and jq

One curl and jq command that lists the last 100 failed jobs, groups them by error.category, code and retryable, and prints a sample job id for each group.

4 min readSume
All posts

List failed jobs with GET /v1/jobs?status=failed&limit=100 and group them with jq by error.category, error.code and error.retryable. Ten groups of failures read faster than one hundred stack of rows, and the retryable flag tells you which groups a retry could fix. The command prints a count, the category, the code, the flag and one sample job id per group, so you can open a single example from each.

The error object

Failed and canceled jobs carry a public-safe error summary. Raw task ids, signed URLs, request bodies and stack traces are left out.

Job error fields from the OpenAPI spec, read 2026-10-08
FieldMeaning
codeStable Sume error code, kept for compatibility (may be absent)
categoryThe broad class of failure
stageWhere in the pipeline it failed
retryableWhether a retry could succeed
retry_after_secondsHow long to wait before a retry
public_reason and next_actionA short reason and the suggested next step

The command

Save it as failed-by-cause.sh. It needs curl and jq. The code field can be missing, so the filter prints a dash instead of null.

#!/usr/bin/env bash
set -euo pipefail
: "${SUME_API_KEY:?set SUME_API_KEY}"
curl -sf -H "x-api-key: $SUME_API_KEY" \
  "https://api.sume.com/v1/jobs?status=failed&limit=100" |
jq -r '.data.jobs
  | group_by([.error.category, .error.code, .error.retryable])
  | map({n: length, cat: .[0].error.category, code: .[0].error.code,
         retry: .[0].error.retryable, sample: .[0].id})
  | sort_by(-.n)[]
  | [.n, .cat, (.code // "-"), .retry, .sample] | @tsv'

Reading the output

Run it after a bad night, not as a monitor. For alerts, count failures from your webhook receiver, where each job.failed event arrives as it happens.

A worked reading: suppose the output has 14 rows in one group with retryable true and a retry_after_seconds value, and 3 rows in another group with retryable false. The retryable flag on the first group says a retry may succeed, so resubmit it slowly with the same keys after the wait. The second group is a request problem, so open the sample id, read public_reason, fix the input, and only then send new work.

  • A large group with retryable true is where a retry with the same Idempotency-Key is worth trying. Wait for retry_after_seconds first.
  • A large group with retryable false needs a fix to the request, such as a bad reference URL. Open the sample job and read public_reason.
  • The list returns at most 100 jobs per page. Pass next_cursor as starting_after to read older pages.
  • Each key reads only the jobs its own member created, so run the command once per key.

Variations

Swap the group key to print one table per stage by using .error.stage in the group_by call. Add a type filter, for example type=video, to the query string to look at a single product. To limit the check to a time window, store the first and last created_at values per group with min and max in jq. Because each of these runs is one read, it is cheap against the read budget, which is far larger than the write budget on every plan.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume