Size a Recast batch from generation_limits, not wave_size_hint

Each Sume submit returns generation_limits and idempotency_hit. A Python loop computes headroom as concurrency minus active minus queued.

4 min readSume
All posts

A Sume submit response for a paid generation carries more than a job id. For video-router calls the envelope's data holds job, idempotency_hit, usage and generation_limits. The last one is a snapshot of your workspace's admission state, and it is the cheapest way to pace a batch without a separate call.

The fields that matter are concurrency_limit (jobs that can be processing), queued_jobs_limit (extra jobs that can wait), active_generation_jobs, queued_generation_jobs and queue_capacity_remaining. There is also a wave_size_hint, and the docs are blunt about it: it is a submission-wave hint, not a concurrency limit, and you should never use it to size in-flight work.

The headroom formula

The generation admission guide gives the budget for new in-flight work as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining. At zero, wait and look again. With a limit of 100, 30 processing and 10 queued, the budget is 60.

Plan defaults for reference

Default processing concurrency and queue capacity by plan, from the generation admission guide (read 2026-10-03)
PlanProcessingQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

The docs tell you to prefer the effective concurrency_limit field over this table, because an admin override can raise it. The loop below does exactly that.

A pacing loop

The script submits Recast jobs one at a time, each with its own Idempotency-Key, and recomputes headroom from the generation_limits on every response. It stops when headroom reaches zero and prints what is left for the next wave. idempotency_hit tells you whether a response was a replay of an earlier submit. The snapshot is taken at submit time, so it can be stale a moment later, and a full queue is a 429 queue_full that you retry with the same key.

import json, os, urllib.request

BASE = "https://api.sume.com/v1"
HEAD = {"x-api-key": os.environ["SUME_API_KEY"], "Content-Type": "application/json"}


def headroom(gl: dict) -> int:
    free = gl["concurrency_limit"] - gl["active_generation_jobs"] - gl["queued_generation_jobs"]
    return max(0, min(free, gl["queue_capacity_remaining"]))  # never wave_size_hint


def submit(order: str) -> dict:
    body = {"model": "h3-max-recast", "mode": "async",
            "video_url": f"https://example.com/{order}.mp4",
            "reference_image_urls": ["https://example.com/host.jpg"]}
    req = urllib.request.Request(BASE + "/video-router/generate", json.dumps(body).encode(),
                                 {**HEAD, "Idempotency-Key": "recast-" + order})
    with urllib.request.urlopen(req, timeout=60) as r:
        return json.load(r)["data"]


pending, room = [f"order-{n}" for n in range(1, 6)], 1
while pending and room > 0:
    data = submit(pending[0])
    room = headroom(data["generation_limits"]) if data.get("generation_limits") else 0
    print(data["job"]["id"], "replay" if data["idempotency_hit"] else "new", "headroom", room)
    pending.pop(0)
print("left for the next wave:", pending)

Rules for a batch

  • Queued is not failed. Sume accepts valid jobs as queued while queue capacity remains, so you can submit past the processing limit.
  • A missing generation_limits is null, so the sample treats it as no headroom and waits to look again.
  • Keep the job ids as you go and poll GET /v1/jobs/{id}/status for the wave. A finished job frees a seat.
  • Reuse the same key for a retry of the same order, so a repeat of the call returns the original job rather than a second one.

The field list, the sizing example and the queue rules are in Generation admission. Status polling is in the jobs guide.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume