Batch Seedance and Kling jobs: generation_limits and wave_size_hint

Submitting many Seedance or Kling clips at once? How Sume's generation_limits fields give a safe in-flight budget, and why wave_size_hint is not concurrency.

5 min readSume
All posts

To submit a batch of Seedance or Kling clips safely, size each wave from generation_limits in the last submit response: new in-flight work is concurrency_limit - active_generation_jobs - queued_generation_jobs, capped by queue_capacity_remaining. wave_size_hint is only a submission-wave hint that includes queue slots. It is not your concurrency and it is not a processing width.

This matters most on video, where a Seedance 2.5 clip can run up to 30 seconds and Kling 3 up to 15, so each job occupies a processing slot for a while. The rules below come from Sume's generation admission page.

What happens when I submit more clips than my plan processes at once?

Concurrency is a dispatch limit, not a submit limit. If your workspace is at its processing limit, Sume still accepts valid jobs as queued while queue capacity remains, and workers move them to processing as slots free up. A queued job is normal, and the right client behavior is to store the job id and poll with backoff.

Only when queue capacity is also full does a submit fail, with 429 queue_full. That is different from 429 rate_limited, which is request-volume protection. Balance is a third control: if Sume cannot reserve the estimated cost, the submit fails with 402 insufficient_credits before any provider work starts.

What are the default limits by plan?

Concurrency is plan-based, and prepaid top-ups do not raise it. Queue capacity defaults to max(3, concurrency_limit x 5). The dashboard Concurrency tab and the concurrency_limit field are the source of truth, because admin overrides and org floors change the numbers.

Default generation limits by plan, read 2026-10-02
PlanProcessingQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Why is wave_size_hint not my concurrency?

The hint is max(1, floor(queue_capacity_remaining x 0.75)). On an empty Pro workspace that is 18, because 24 accepted slots times 0.75. But only 4 of those jobs can process at a time. Sume's docs say plainly never to use the hint to size in-flight work, and never to present it as concurrency.

The counts are also a snapshot. They can change the moment after the response as workers claim jobs or other clients submit. Treat them as a conservative guide, and count every job you submit against your budget until you read a fresh snapshot.

How do I turn the snapshot into a batch loop?

The Video Router submit response carries generation_limits when Sume can compute it. The documented /v1/videos submit body lists only id, polling_url, status and model, so this sketch uses POST /v1/video-router/generate and reads the field from either the top level or a data envelope. It stops adding work at zero headroom.

import os, time, uuid, requests

URL = "https://api.sume.com/v1/video-router/generate"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def headroom(gl: dict) -> int:
    free = gl["concurrency_limit"] - gl["active_generation_jobs"] \
        - gl["queued_generation_jobs"]
    return max(0, min(free, gl["queue_capacity_remaining"]))

def submit(prompt: str):
    body = {"model": "seedance-2.5", "prompt": prompt,
            "resolution": "480p", "duration": 4, "mode": "async"}
    r = requests.post(URL, json=body, timeout=60,
                      headers={**H, "Idempotency-Key": str(uuid.uuid4())})
    r.raise_for_status()
    j = r.json()
    return (j.get("generation_limits")
            or j.get("data", {}).get("generation_limits"))

if __name__ == "__main__":
    gl = submit("A paper boat drifting down a rain gutter")
    while gl and headroom(gl) > 0:
        gl = submit("A paper boat drifting down a rain gutter")
    print("stop: no headroom", gl)

What does Sume not give me for batches?

There is no per-job queue position or ETA, only queue counts and remaining accepted capacity. Queue expiration and a client-supplied fail-fast queue length are not public options. If you hit queue_full, stop adding work, poll existing jobs until one is terminal, cancel queued jobs you no longer need (cancel works only before generation starts, and later returns 409 job_generation_already_started), and retry with the same idempotency key.

Use the cheapest test settings while you tune the loop. The sample uses 480p and 4 seconds, the lowest resolution and the minimum duration seedance-2.5 lists. See Video Router for per-model limits and Errors and rate limits for the retry table.

What should my loop do when headroom hits zero?

Stop submitting and switch to polling. Each terminal job frees a slot, and the next submit response brings a fresh generation_limits snapshot to read. Re-derive headroom from that snapshot, not from your own counters alone, because other clients on the same workspace share the same limits.

Keep the job ids in durable storage. If your process dies mid-batch, the jobs keep running on Sume, and you can resume polling by id instead of resubmitting.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume