GitHub Actions matrix cap is 256 jobs: use Sume bulk runs instead
A GitHub Actions matrix can create 256 jobs, and a Sume bulk run takes 100 items. How to split a batch between them so concurrency and queue limits hold.

Do not fan a Sume batch out into a GitHub Actions matrix. One job can submit a whole bulk run of up to 100 items, and the matrix is capped at 256 jobs per workflow run, so a matrix wastes runner minutes and still collides with Sume's workspace concurrency.
The Actions limits page (read 2026-10-10) says a job matrix can generate a maximum of 256 jobs per workflow run. Concurrent job limits by plan are 20 for Free, 40 for Pro, 60 for Team and 500 for Enterprise. GitHub-hosted jobs run up to 6 hours, self-hosted up to 5 days, and a workflow run is cancelled at 35 days.
Why the matrix is the wrong shape
A matrix of 200 jobs, each submitting one Sume generation, creates 200 runners that mostly wait. Your Actions plan allows 20 to 500 concurrent jobs, while the Sume Free plan admits 6 jobs in total and Pro admits 24. The rest of the requests come back 429 queue_full, and every one of those runner jobs has to retry. You pay for runner time to discover a limit that is in the docs.
Sume's bulk runs take 1 to 100 items and a required concurrency window of 1 to 16. The API keeps that many child Format runs in flight and queues the rest on its side, so one workflow step submits the lot. Workspace generation concurrency still applies to the child runs, so the window is a ceiling, not a guarantee.
| Approach | Workflow jobs | Sume submits | Notes |
|---|---|---|---|
| One job per item | 250 of 256 allowed | 250 single runs | Hits plan concurrency and queue_full |
| Matrix of 3 chunks | 3 | 3 bulk runs (100, 100, 50) | Each chunk waits on its own window |
| One job, sequential | 1 | 3 bulk runs | Fewest runner minutes |
A workable layout
Split the list into chunks of at most 100 in a setup step. Give each chunk its own Idempotency-Key, for example the commit sha plus the chunk number, so a re-run of the workflow does not spend twice. The Sume docs say to mint a fresh key per batch, and that replaying a spent key returns 202 with the earlier receipt rather than starting new work.
Then poll the queue resource, GET /v1/format-run-queues/{id}, from the same job, or let a webhook call a small endpoint that finishes the pipeline. A hosted job can run for 6 hours, which is more than a normal bulk run needs, but the 35 day workflow ceiling matters only if you park a workflow on a human approval.
- Chunk size 100 or less.
concurrencyat or below your plan's concurrency.- Idempotency key from sha plus chunk number.
- Use a webhook instead of a long poll when runs take many minutes.
Rate limits to remember
The Actions page also lists a GITHUB_TOKEN limit of 1,000 requests per hour per repository. That is a GitHub-side budget for API calls to GitHub, for example posting results as comments, so batch your result posts. On the Sume side, read and write limits are per plan and a 429 carries retry-after, so a polling loop should sleep for that value and not spin.
Sources
Related posts
More in Integrations
- GitHub environment reviewers: approve before a paid Sume run
Put the Sume key in a GitHub environment with required reviewers, so a workflow pauses for approval before it starts a paid run, then cap the run's spend.
- GitLab resource_group: one Sume bulk queue at a time
Use a GitLab resource_group so two pipelines never submit a Sume bulk run together, and pick the process mode that matches how stale a queued batch may be.
- Cloud Scheduler retryCount max 5: key the Sume submit by slot
Cloud Scheduler retries a failed target up to 5 times, and a retried call looks like a new request. Key the Sume submit by the schedule slot, not the retry.
- Grafana alert webhook with HMAC to a Sume run: incident explainer
Grafana's webhook contact point can sign alerts with HMAC over timestamp:body. Verify it, then start a Sume Format run per firing alert with a stable key.
Written by Sume