GitHub Actions concurrency queue: max vs single, and Sume queue_full

GitHub concurrency keeps one pending run by default, up to 100 with queue: max. See what that means for Sume submits and why cancel-in-progress is risky.

5 min readSume
All posts

A GitHub Actions concurrency group lets one run proceed while others wait, and by default it keeps at most one waiting run, replacing any older pending one. If every push should produce a Sume run, that default silently drops the middle ones, so you need queue: max, which GitHub documents as allowing up to 100 pending runs per group. Sume has its own separate queue and its own 429 queue_full, and the two should not be confused.

GitHub's behaviour is from Control the concurrency of workflows and jobs, read 2026-10-10. Sume's admission rules are from Generation admission.

What does each GitHub setting do?

A group name can be any string or expression, using contexts such as github, inputs, vars, needs, strategy and matrix, and group names are case insensitive. The concurrency key works at workflow level or on a single job. With the default queue: single, at most one run can be pending, and a newer arrival cancels and replaces the older pending one. With queue: max, up to 100 can be pending and they run in the order they started waiting. GitHub also notes that queue: max and cancel-in-progress: true cannot be combined because they describe conflicting behaviours.

GitHub concurrency settings (read 2026-10-10)
SettingEffect on runs in one groupFit for a Sume submit job
Default (queue: single)One pending run; a newer one replaces the older pending runFine for a render of the latest commit only
queue: maxUp to 100 pending runs, first in first outFine when each run is its own paid item
cancel-in-progress: trueTerminates the running job when a new one arrivesDangerous after the paid request was sent
Group expressionAny string built from contextsUse the SKU or branch so unrelated work does not wait

Why is cancel-in-progress the risky one?

Canceling the GitHub job stops your runner, not the Sume work. Once the submit request returned a run id, the render continues and bills on Sume's side. Create a run and Runs and results describe the pieces you control: POST /v1/format-runs/{run_id}/cancel stops a run in progress, and the call is idempotent, returning cancel_effect of canceled or no_op. You pay for generation completed before the cancel.

If a newer push should replace an older render, make the cancel an explicit step with its own condition instead of hoping the runner cancellation did it. If you only want to skip stale work before it is paid for, keep cancel-in-progress off the submit job and put it on the cheap lint job.

Where does Sume's queue come in?

Sume admits paid jobs queue-first. Generation concurrency limits processing, not submission: a valid job is accepted as queued while queue capacity remains. The plan table on the admission page gives processing concurrency of 1 (Free), 4 (Pro), 8 (Startup) and 20 (Scale and Enterprise), with default queue capacity of 5, 20, 40 and 100. A submit fails with 429 queue_full only when concurrency and queue are both full, and 429 rate_limited is a different, request-volume limit.

A GitHub group that lets 100 runs wait does not mean 100 Sume jobs at once. Each waiting run starts its own submit later, so GitHub's queue naturally paces you. If several groups can submit in parallel, read generation_limits.queue_capacity_remaining from a submit response, and treat the dashboard concurrency figure as the source of truth rather than a copied table.

A safe configuration

Use one group per thing you are producing, queue: max so no item is dropped, no cancel-in-progress on the submit job, and an Idempotency-Key derived from the item and a version you bump on purpose. On 429 queue_full, wait and retry with the same key; the admission page says a retry with the same key is the correct move.

Remember the limit on the GitHub side too: 100 pending is a ceiling. If a nightly job can create more than that, submit through a Sume bulk queue instead, which takes 1 to 100 items with a concurrency window of 1 to 16 on the server.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume