GitHub Actions weekly Sume bulk queue that fails on counts.failed

A scheduled workflow that creates a Sume bulk queue, polls it through transient 429 and 503, and turns a completed queue with failed items into a red job.

5 min readSume
All posts

To run a weekly Sume bulk queue from GitHub Actions, create the queue in one step, poll GET /v1/format-run-queues/{id} until status is completed, and then fail the job if counts.failed or counts.canceled is above zero. The last check is the part people skip. A queue reaches completed when every item is terminal, so a green poll loop on its own tells you nothing about whether the videos exist.

The workflow below keeps the batch in an items.json file in the repository. It is an array of item bodies, each the same shape as a single Format run.

The workflow

The idempotency key is the ISO week, so re-running the workflow in the same week replays the same queue instead of starting a second batch.

name: weekly-sume-batch
on:
  schedule: [{ cron: "0 9 * * 1" }]
  workflow_dispatch:
jobs:
  batch:
    runs-on: ubuntu-latest
    timeout-minutes: 120
    env:
      SUME_API_KEY: ${{ secrets.SUME_API_KEY }}
      API: https://api.sume.com/v1
      FORMAT: acme/product-promo
    steps:
      - uses: actions/checkout@v4
      - shell: bash
        run: |
          set -euo pipefail
          KEY="weekly-$(date -u +%G-W%V)"
          Q=$(jq -c '{concurrency: 3, items: .}' items.json | curl -sS --fail-with-body \
            -X POST "$API/formats/$FORMAT/bulk-runs" -H "Authorization: Bearer $SUME_API_KEY" \
            -H "Content-Type: application/json" -H "Idempotency-Key: $KEY" -d @-)
          ID=$(jq -r .data.id <<<"$Q"); echo "queue $ID"
          while :; do
            R=$(curl -sS --fail-with-body "$API/format-run-queues/$ID" \
              -H "Authorization: Bearer $SUME_API_KEY") || { sleep 60; continue; }
            [ "$(jq -r .data.status <<<"$R")" = completed ] && break
            sleep 60
          done
          jq -c .data.counts <<<"$R"
          jq -e '.data.counts.failed == 0 and .data.counts.canceled == 0' <<<"$R" >/dev/null

What each part is doing

The create call needs a key with formats:write, and the poll needs formats:read. Store the key as a repository secret. A key made before the Format API shipped has neither scope and fails with 403 insufficient_scope; make a new key instead of retrying.

The || { sleep 60; continue; } branch exists because a 429 or 503 during a poll is transient. The queue keeps draining while your poller waits, and the docs say not to read those codes as a failed queue. The same branch also swallows permanent errors such as a revoked key, so the job-level timeout-minutes is the ceiling. Hitting it stops only the poller. The queue itself is a server-side list and keeps running, which is why the step prints the queue id on its own line.

What the final step tells you (read 2026-10-07)
Queue state`counts` showsJob result
completed, all items donefailed: 0, canceled: 0Green
completed, some items failedfailed: NRed; read each failed item's run receipt for the cause
completed, a child was canceledcanceled: NRed; someone stopped a run by hand
Still running at the timeoutqueued and running above 0Red from the timeout; the queue is not canceled

When one job is too short for the batch

Video work is slow. The docs describe long-form runs as 15 to 30 minutes each, and if each item is that long, a queue with concurrency: 3 and 100 items runs about 34 waves of it, and a single two-hour job will not outlast that. Two patterns work. Either create the queue in one workflow, save the queue id, and poll in a later scheduled run that reads it, or raise concurrency (up to 16), which is still bounded by the generation concurrency your plan allows.

Either way the poller is stateless: it needs only the queue id and a key with formats:read. Polls come out of the read budget, which is separate from the write budget and forty times larger, so a 60-second sleep costs nothing meaningful.

Things worth knowing before you copy it

The queue has no webhook, and the API has no list-queues endpoint, so the log line with the queue id is your only handle if the job dies. If you lose it, send the same create call again with the same key and body: the replay returns 202 and the original queue.

Set generation_spend_cap_usd on each item in items.json. The queue has no cap of its own, so the worst case for a run of this workflow is the sum of the item caps. See worst-case spend of a 100-item queue.

Do not add on_active_run to the items. The bulk controller starts every item with allow, so skip and reject are ignored inside a queue.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume