Submit 60 Gemini Omni 4K jobs on a Pro plan: pace in waves of 24
A Pro workspace holds 24 accepted generation jobs (4 running, 20 queued). Pace 60 Omni 4K submits in waves with asyncio and retry queue_full safely.

On a Sume Pro workspace you can hold 24 paid generation jobs at once, 4 processing and 20 queued, so 60 Gemini Omni Flash 1.1 jobs at 4K need three waves: 24, 24 and 12. Submit up to the accepted limit, let jobs finish, then top up as slots free, and treat 429 queue_full as a signal to wait, not an error to count.
The numbers that decide the wave size
Sume separates request rate from generation capacity. Concurrency is the number of jobs running. Queue capacity is how many more can wait. The default queue is max(3, concurrency times 5), and the accepted total is concurrency plus queue. Full concurrency alone is not an error. Only when the queue is also full does a submit fail with 429 queue_full.
Omni at 4K is 3 to 10 seconds per clip, so each job is short but still counts against the same slots as a 30-second Seedance render. Slots are per job, not per second of video.
One more reason to pace: money. Sume reserves the estimated cost of each job when it accepts the request, so 24 accepted 4K jobs hold 24 reservations at once. Pacing to your accepted limit also paces your reserved balance, which is useful if a 402 insufficient_credits would otherwise appear halfway through a run and leave a batch half done.
| Plan | Processing | Queue | Accepted jobs | Waves for 60 clips |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | 10 |
| Pro | 4 | 20 | 24 | 3 |
| Startup | 8 | 40 | 48 | 2 |
| Scale | 20 | 100 | 120 | 1 |
A wave runner
The script below keeps at most CAP jobs open at any time with an asyncio semaphore. It sends each submit with a stable Idempotency-Key, so a retry after queue_full cannot create a duplicate, and it polls until terminal before releasing the slot. It uses asyncio.to_thread around urllib so it needs no dependencies. Set CAP to your plan's accepted number, or lower it if you share the workspace.
Here is the arithmetic in a sentence. 60 jobs at 24 accepted per wave is 2.5 waves, which rounds up to 3, and the last wave holds 12 jobs, so the final wave leaves half the capacity idle. If you can add 12 more jobs to that run, the wave is full and the whole run takes the same three rounds.
import asyncio, json, os, urllib.request, urllib.error
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"}
CAP = 24 # Pro: 4 processing + 20 queued
def call(method, path, body=None, extra=None):
req = urllib.request.Request("https://api.sume.com" + path, method=method,
data=json.dumps(body).encode() if body else None, headers={**H, **(extra or {})})
try:
with urllib.request.urlopen(req, timeout=30) as r:
return r.status, json.loads(r.read())
except urllib.error.HTTPError as e:
return e.code, json.loads(e.read() or b"{}")
async def one(i, sem):
body = {"model": "gemini-omni-flash-1.1", "prompt": f"Product turntable, take {i}",
"duration": 8, "resolution": "4K", "aspect_ratio": "16:9"}
key = {"Idempotency-Key": f"omni4k-{i}"}
async with sem:
code, job = await asyncio.to_thread(call, "POST", "/v1/videos", body, key)
while code != 202: # queue_full or rate_limited: wait, resend with the same key
await asyncio.sleep(20)
code, job = await asyncio.to_thread(call, "POST", "/v1/videos", body, key)
while True:
_, st = await asyncio.to_thread(call, "GET", "/v1/videos/" + job["id"])
if st.get("status") in ("completed", "failed", "cancelled"):
return i, st["status"]
await asyncio.sleep(10)
async def main():
sem = asyncio.Semaphore(CAP)
print(await asyncio.gather(*(one(i, sem) for i in range(60))))
asyncio.run(main())Why the semaphore sits around the whole job
Holding the slot from submit to terminal is what keeps you under the accepted limit. If you release it after the submit, you can enqueue all 60 in seconds and the 25th gets queue_full. The Sume docs describe the same pacing rule: budget new in-flight work from the live counts, and refresh before you widen.
Do not use the wave_size_hint as a concurrency number. The admission page defines it as a submission hint, max of 1 and 75 percent of remaining queue capacity, and says never to use it to size in-flight work.
Retry rules for the loop
Test the pacer with CAP set to 2 and a range of 4 before the real run. You should see two submits, a poll wait, then the next two. If you see all four submits at once, the semaphore is in the wrong place. A cheap dry run with a low resolution costs far less than learning the same thing at 4K.
- On 429 queue_full or rate_limited, wait and resend with the same Idempotency-Key. The reservation for the failed admission is released.
- On 402 insufficient_credits, stop the whole run. Waiting will not fix a balance.
- On 400 unsupported_capability, fix the request. Omni accepts 3 to 10 seconds, so a duration of 12 is rejected before any provider work.
- Cancel queued jobs you no longer want. Cancel works only before generation starts.
Sources
Related posts
More in Developers
- Perl HTTP::Tiny: submit a MiniMax H3 video job and save it
A 23-line Perl script with core modules only: POST /v1/videos for minimax-h3, poll, download. A 5-second 768p clip is $0.375; 480p is $0.3125.
- PHP cURL: save a Sume-generated image to disk with CURLOPT_FILE
Two cURL calls in plain PHP: POST to /v1/images, read data[0].url on a 200, then stream the file to disk with CURLOPT_FILE. Handles 202 and a missing key.
- PHP cURL: submit a Seedance 2.5 video job, poll it, save the MP4
PHP with only ext-curl: POST /v1/videos for seedance-2.5, poll polling_url until completed, then download content?index=0; a 5-second 480p test costs $1.34.
- Pick a Sume video model by script: filter /v1/videos/models
Instead of guessing, call GET /v1/videos/models and filter by ratio, resolution and last frame. A 15-line Python script lists the Sume video models that fit.
Written by Sume