Batch Seedance and Kling jobs: generation_limits and wave_size_hint
Submitting many Seedance or Kling clips at once? How Sume's generation_limits fields give a safe in-flight budget, and why wave_size_hint is not concurrency.

To submit a batch of Seedance or Kling clips safely, size each wave from generation_limits in the last submit response: new in-flight work is concurrency_limit - active_generation_jobs - queued_generation_jobs, capped by queue_capacity_remaining. wave_size_hint is only a submission-wave hint that includes queue slots. It is not your concurrency and it is not a processing width.
This matters most on video, where a Seedance 2.5 clip can run up to 30 seconds and Kling 3 up to 15, so each job occupies a processing slot for a while. The rules below come from Sume's generation admission page.
What happens when I submit more clips than my plan processes at once?
Concurrency is a dispatch limit, not a submit limit. If your workspace is at its processing limit, Sume still accepts valid jobs as queued while queue capacity remains, and workers move them to processing as slots free up. A queued job is normal, and the right client behavior is to store the job id and poll with backoff.
Only when queue capacity is also full does a submit fail, with 429 queue_full. That is different from 429 rate_limited, which is request-volume protection. Balance is a third control: if Sume cannot reserve the estimated cost, the submit fails with 402 insufficient_credits before any provider work starts.
What are the default limits by plan?
Concurrency is plan-based, and prepaid top-ups do not raise it. Queue capacity defaults to max(3, concurrency_limit x 5). The dashboard Concurrency tab and the concurrency_limit field are the source of truth, because admin overrides and org floors change the numbers.
| Plan | Processing | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Why is wave_size_hint not my concurrency?
The hint is max(1, floor(queue_capacity_remaining x 0.75)). On an empty Pro workspace that is 18, because 24 accepted slots times 0.75. But only 4 of those jobs can process at a time. Sume's docs say plainly never to use the hint to size in-flight work, and never to present it as concurrency.
The counts are also a snapshot. They can change the moment after the response as workers claim jobs or other clients submit. Treat them as a conservative guide, and count every job you submit against your budget until you read a fresh snapshot.
How do I turn the snapshot into a batch loop?
The Video Router submit response carries generation_limits when Sume can compute it. The documented /v1/videos submit body lists only id, polling_url, status and model, so this sketch uses POST /v1/video-router/generate and reads the field from either the top level or a data envelope. It stops adding work at zero headroom.
import os, time, uuid, requests
URL = "https://api.sume.com/v1/video-router/generate"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def headroom(gl: dict) -> int:
free = gl["concurrency_limit"] - gl["active_generation_jobs"] \
- gl["queued_generation_jobs"]
return max(0, min(free, gl["queue_capacity_remaining"]))
def submit(prompt: str):
body = {"model": "seedance-2.5", "prompt": prompt,
"resolution": "480p", "duration": 4, "mode": "async"}
r = requests.post(URL, json=body, timeout=60,
headers={**H, "Idempotency-Key": str(uuid.uuid4())})
r.raise_for_status()
j = r.json()
return (j.get("generation_limits")
or j.get("data", {}).get("generation_limits"))
if __name__ == "__main__":
gl = submit("A paper boat drifting down a rain gutter")
while gl and headroom(gl) > 0:
gl = submit("A paper boat drifting down a rain gutter")
print("stop: no headroom", gl)
What does Sume not give me for batches?
There is no per-job queue position or ETA, only queue counts and remaining accepted capacity. Queue expiration and a client-supplied fail-fast queue length are not public options. If you hit queue_full, stop adding work, poll existing jobs until one is terminal, cancel queued jobs you no longer need (cancel works only before generation starts, and later returns 409 job_generation_already_started), and retry with the same idempotency key.
Use the cheapest test settings while you tune the loop. The sample uses 480p and 4 seconds, the lowest resolution and the minimum duration seedance-2.5 lists. See Video Router for per-model limits and Errors and rate limits for the retry table.
What should my loop do when headroom hits zero?
Stop submitting and switch to polling. Each terminal job frees a slot, and the next submit response brings a fresh generation_limits snapshot to read. Re-derive headroom from that snapshot, not from your own counters alone, because other clients on the same workspace share the same limits.
Keep the job ids in durable storage. If your process dies mid-batch, the jobs keep running on Sume, and you can resume polling by id instead of resubmitting.
Sources
Related posts
More in Developers
- Seedance or Kling job failed: could not download an input media URL
A Seedance or Kling job can fail with 'Could not download an input media URL'. Why Sume rejects a frame or reference URL, and a preflight check for it.
- Seedance and Kling submit errors: which to retry and which to fix
A Sume video submit can fail with 402, 429 queue_full, 429 rate_limited or 400. Which to retry with the same key and which need a changed request.
- Seedance or Kling video 409: job_not_completed vs job_failed
A 409 from GET /v1/videos/{id}/content means two opposite things on Sume: job_not_completed is retryable, job_failed is not. How to tell them apart.
- One image to Seedance: reference, or first frame?
On Sume a single reference image with no frame field is priced and routed as reference-to-video; add a first frame to get image-to-video. How to choose.
Written by Sume