100 Omni clips on Sume Pro: queue_full after 24 jobs

Sume Pro runs 4 jobs at once and accepts 24 in total. Submitting 100 Omni clips at once returns 429 queue_full. A Python wave loop that retries safely.

6 min readSume
All posts

If you send 100 Gemini Omni Flash 1.1 jobs to Sume on a Pro plan in one burst, about 24 are accepted and the rest fail with 429 queue_full. Pro processes 4 generation jobs at once and queues 20 more, so accepted capacity is 24. Submit in waves, reuse one Idempotency-Key per clip, and retry after the wait. The numbers come from Sume's generation-admission docs.

A full processing slot is not an error: Sume accepts valid jobs as queued while queue room remains. The error appears only when both processing and the queue are full.

Capacity by plan

Concurrency is set by plan, and prepaid top-ups do not raise it. The dashboard Concurrency tab and the generation_limits object in each submit response show the effective numbers.

Sume generation concurrency and queue by plan, docs read 2026-10-08
PlanProcessing at onceQueueAccepted in totalWrites per minute
Free156120
Pro42024300
Startup84048600
Scale201001201200

What 100 clips cost, and what Sume holds

At 8 seconds and 720p, an Omni clip bills $1.00 on Sume ($0.125 per second, provider list $0.10 times 1.25). A hundred of them is $100.00. At submit Sume reserves the estimate from the workspace balance, and a request that cannot be reserved fails with 402 insufficient_credits before any provider work starts. Google's direct price for the same model is about $0.10 per second at 720p on its pricing page, read 2026-10-08, but that is a separate account and limit system.

Pick the resolution before you queue, because a 4K clip is $0.375 per second on Sume, three times the 720p rate. A 360p draft pass is $0.0375 per second, so 100 drafts of 8 seconds cost $30.00, and you can approve a shortlist before the 720p or 1080p finals. Each hundred-clip pass is also a chance to learn your real queue time, since Pro starts only 4 at once and the other 20 wait.

A wave loop that is safe to restart

The key per clip means a crash and rerun does not create duplicate jobs: Sume says a replay returns the original job. On 429, the loop waits for retry-after when present. It sets 30 seconds otherwise.

import os, time, requests

API = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(i, prompt):
    body = {"model": "gemini-omni-flash-1.1", "prompt": prompt,
            "duration": 8, "resolution": "720p", "aspect_ratio": "9:16"}
    headers = {**AUTH, "Idempotency-Key": f"omni-wave-{i:03d}"}
    while True:
        r = requests.post(API, json=body, headers=headers, timeout=60)
        if r.status_code == 429:
            time.sleep(int(r.headers.get("retry-after", "30")))
            continue
        r.raise_for_status()
        return r.json()["id"]

if __name__ == "__main__":
    ids = [submit(i, f"Product hook number {i}, handheld, daylight")
           for i in range(100)]
    print(len(ids), "jobs accepted")

Poll cheaply

Do not spend the write budget on status checks. A poll is a read, and reads have their own, much larger bucket. Poll GET /v1/jobs/:id/status with backoff, or give each job a webhook. Do not retry unsafe submits without the key; Sume's errors page says so.

Cost of the full run

One hundred 10-second clips at 720p cost $125.00 on Sume ($1.25 each), at 1080p $187.50 and at 360p $37.50. Draft the whole set at 360p only if the prompts vary a lot; otherwise draft a sample of ten and run the rest at the final resolution.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume