Count queued and processing jobs with GET /v1/jobs before a wave

Page GET /v1/jobs with status=queued and status=processing, subtract from concurrency_limit, and submit only that many Seedance 2.5 or Omni clips.

4 min readSume
All posts

To know how many video jobs are in flight, call GET /v1/jobs twice, once with status=queued and once with status=processing, and follow data.next_cursor until it disappears. Subtract both counts from your workspace's concurrency_limit, and submit at most that many new clips.

This works after a restart, when you no longer hold the generation_limits block from your last submit response.

Which numbers does the headroom formula use?

The admission docs size new in-flight work as max(0, concurrency_limit - active - queued), capped by queue_capacity_remaining. The active and queued counts are what the list gives you; concurrency_limit comes from the generation_limits object on a submit response or from the dashboard Concurrency tab.

Do not use wave_size_hint as a width. The docs define it as a submission hint that includes queue slots, not as concurrency.

Fields used for headroom (Sume docs and OpenAPI, read 2026-10-09)
FieldWhere it comes fromRole
concurrency_limitgeneration_limits on a submit envelopeEffective processing cap
status=processingGET /v1/jobs filterActive jobs
status=queuedGET /v1/jobs filterWaiting jobs
queue_capacity_remaininggeneration_limits on a submit envelopeUpper bound before queue_full

What does the Python look like?

Pages hold up to 100 jobs, newest first. The cursor is opaque, and the last page simply has no next_cursor, so loop on its absence instead of on a short page.

import os, requests

BASE = "https://api.sume.com/v1"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def count(status):
    n, cursor = 0, None
    while True:
        q = {"status": status, "limit": 100}
        if cursor:
            q["starting_after"] = cursor
        r = requests.get(f"{BASE}/jobs", headers=HEAD, params=q, timeout=30)
        r.raise_for_status()
        data = r.json()["data"]
        n += len(data["jobs"])
        cursor = data.get("next_cursor")
        if not cursor:
            return n

limit = int(os.environ.get("SUME_CONCURRENCY", "4"))
room = max(0, limit - count("processing") - count("queued"))
print("submit at most", room, "new clips")

Why is the list only a lower bound?

An API key reads the jobs its own member created. If teammates submit into the same workspace with other keys, their jobs count against the workspace cap but never appear in your list. Treat your count as a floor, and refresh from a live generation_limits block whenever you have one.

The status filter accepts queued, processing, completed, failed and canceled. Check that a filter name is spelled as in the OpenAPI before you rely on the count.

When should you re-count?

Counts are a snapshot. Workers claim queued jobs and other clients submit between your read and your write. Re-count before each wave, count each job you just submitted against the headroom until the next snapshot, and keep reading the list at a modest pace, since list calls spend the read budget rather than the write budget.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume