Count queued and processing jobs with GET /v1/jobs before a wave
Page GET /v1/jobs with status=queued and status=processing, subtract from concurrency_limit, and submit only that many Seedance 2.5 or Omni clips.

To know how many video jobs are in flight, call GET /v1/jobs twice, once with status=queued and once with status=processing, and follow data.next_cursor until it disappears. Subtract both counts from your workspace's concurrency_limit, and submit at most that many new clips.
This works after a restart, when you no longer hold the generation_limits block from your last submit response.
Which numbers does the headroom formula use?
The admission docs size new in-flight work as max(0, concurrency_limit - active - queued), capped by queue_capacity_remaining. The active and queued counts are what the list gives you; concurrency_limit comes from the generation_limits object on a submit response or from the dashboard Concurrency tab.
Do not use wave_size_hint as a width. The docs define it as a submission hint that includes queue slots, not as concurrency.
| Field | Where it comes from | Role |
|---|---|---|
| concurrency_limit | generation_limits on a submit envelope | Effective processing cap |
| status=processing | GET /v1/jobs filter | Active jobs |
| status=queued | GET /v1/jobs filter | Waiting jobs |
| queue_capacity_remaining | generation_limits on a submit envelope | Upper bound before queue_full |
What does the Python look like?
Pages hold up to 100 jobs, newest first. The cursor is opaque, and the last page simply has no next_cursor, so loop on its absence instead of on a short page.
import os, requests
BASE = "https://api.sume.com/v1"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def count(status):
n, cursor = 0, None
while True:
q = {"status": status, "limit": 100}
if cursor:
q["starting_after"] = cursor
r = requests.get(f"{BASE}/jobs", headers=HEAD, params=q, timeout=30)
r.raise_for_status()
data = r.json()["data"]
n += len(data["jobs"])
cursor = data.get("next_cursor")
if not cursor:
return n
limit = int(os.environ.get("SUME_CONCURRENCY", "4"))
room = max(0, limit - count("processing") - count("queued"))
print("submit at most", room, "new clips")Why is the list only a lower bound?
An API key reads the jobs its own member created. If teammates submit into the same workspace with other keys, their jobs count against the workspace cap but never appear in your list. Treat your count as a floor, and refresh from a live generation_limits block whenever you have one.
The status filter accepts queued, processing, completed, failed and canceled. Check that a filter name is spelled as in the OpenAPI before you rely on the count.
When should you re-count?
Counts are a snapshot. Workers claim queued jobs and other clients submit between your read and your write. Re-count before each wave, count each job you just submitted against the headroom until the next snapshot, and keep reading the list at a modest pace, since list calls spend the read budget rather than the write budget.
Sources
Related posts
More in Developers
- createSumeClient sends x-api-key only: an extra header gives 401
The Sume API accepts Bearer or x-api-key but rejects both at once with 401. A fetch wrapper for createSumeClient that drops the extra header, tested offline.
- curl -w http_code: branch a Sume video submit in bash on 202, 402, 429
A bash submit that saves the body, reads the HTTP code with curl -w and branches on 202, 402, 429 and 5xx. Three Wan 3.0 payloads cost $1.875, $3.75 and $7.50.
- Cut dead air before the first word: STT start time, $0.01 split
Read words[0].start from a Sume STT job, then split the recording from that second for $0.01. For a 3-minute take the whole fix is $0.04. A copyable request.
- Deno fetch: 25 s Wan 3.0 clip costs $3.125, 20-minute deadline
Deno script: submit a 25-second Wan 3.0 clip to Sume at 720p ($3.125), poll with backoff, stop at a 20-minute deadline. Run with deno run -A.
Written by Sume