300 holiday UGC clips on one plan: pace around queue_full 429s
A Pro workspace holds 4 processing and 20 queued jobs. Submit 300 avatar clips, treat queued as normal, and retry queue_full with the same idempotency key.

You can submit far more clips than your workspace processes at once, but only up to its accepted-job capacity: on the Pro plan that is 4 processing plus 20 queued, so 24 paid generation jobs at a time. The 25th returns 429 queue_full, and the documented response is to wait for jobs to finish, then retry with the same idempotency key. A 300-clip UGC batch is therefore a paced submit loop, not one burst.
Treat queued as normal. The docs say concurrency is a dispatch limit, not a submit limit.
What are the limits for each plan?
Generation concurrency is plan-only; prepaid top-ups do not raise it. Queue capacity defaults to max(3, concurrency_limit x 5), and accepted capacity is the sum of the two. Org workspaces have a floor of 10 and Enterprise uses admin overrides, so read the effective generation_limits.concurrency_limit in your responses or the dashboard Concurrency tab rather than trusting this table.
For 300 clips, accepted capacity divided into the batch gives the minimum number of waves. On Pro that is 300 / 24, about 13 waves. How long each wave takes depends on the clips; the docs give no per-job duration, so we do not estimate one.
| Plan | Processing | Queue capacity | Accepted at once | Waves for 300 clips |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | 50 |
| Pro | 4 | 20 | 24 | 13 |
| Startup | 8 | 40 | 48 | 7 |
| Scale | 20 | 100 | 120 | 3 |
Which errors mean wait, and which mean stop?
Four codes show up in a peak batch and they need different handling. queue_full and rate_limited are both 429, but the first means capacity and the second means request volume. insufficient_credits is a 402 returned before provider work starts, because Sume reserves the estimated cost at submit; retrying will not fix it.
A failed admission releases or refunds its reservation where applicable. Do not resubmit a paid request just because your own process timed out; look up the job first. That is why every submit below carries a stable Idempotency-Key per clip.
queue_full(429): stop adding work, poll or cancel, retry with the same key after capacity opens.rate_limited(429): back off usingretry-afterwhen present.insufficient_credits(402): the reservation could not be funded; lower the request cost or fund the workspace per the docs, and do not loop.idempotency_conflict(409): the key was reused for a different payload.
What does the paced submit loop look like?
This script submits 300 avatar clips with a per-clip key and sleeps on queue_full. It uses only the standard library and runs as is with SUME_API_KEY set; the avatar handle is the one used in the Sume docs examples, so replace it with your own ready avatar. It does not poll results, which you should do through job status or webhooks.
Because every clip has its own key, re-running the whole script after a crash is safe: a replayed key returns the original job instead of charging again.
import json, os, time, urllib.error, urllib.request
KEY = os.environ["SUME_API_KEY"]
URL = "https://api.sume.com/v1/avatar-1.0/talking-video"
def submit(n):
body = {"avatar_handle": "sume_clawra", "aspect_ratio": "9:16",
"script": f"Gift idea number {n}: this one arrives wrapped, ships fast, and fits any budget this holiday season."}
req = urllib.request.Request(URL, json.dumps(body).encode(), {
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"Idempotency-Key": f"peak-clip-{n}"})
while True:
try:
with urllib.request.urlopen(req) as r:
return json.load(r)
except urllib.error.HTTPError as e:
err = json.load(e).get("error", {})
if e.code == 429 and err.get("code") == "queue_full":
time.sleep(int(e.headers.get("retry-after", 30)))
continue
raise
for n in range(1, 301):
submit(n)What should you not use `wave_size_hint` for?
Submit responses can include a generation_limits snapshot with a wave_size_hint, defined as max(1, floor(queue_capacity_remaining x 0.75)). On an idle Pro workspace that is 18. The docs are explicit that it is a submission-wave hint only: not a concurrency limit and not a processing width, so never size in-flight work with it.
If a sale moment changes and you no longer need the tail of the batch, cancel jobs that are still queued. Once generation has started, cancel returns 409 job_generation_already_started and the job finishes or fails normally.
Sources
Related posts
More in Use cases
- Hour-long podcast video to clips: the 1800-second source cap
Sume's trim, detach and inspect tools read sources up to 1800 seconds, so a 60-minute episode must be split before import. Limits and a safe split plan.
- How many clips for a 10-minute faceless video? Timeline slot math
A 10-minute faceless video fits Timeline 1.0 comfortably: up to 200 slots, a render near $1.00, and limits on fades and single-pass renders to plan around.
- Instagram Reels AI translation: Korean and Japanese added July 2026
Meta's July 14, 2026 update adds French, German, Italian, Japanese and Korean to Reels translation on Instagram. Here is the full language list and the dates.
- Korean captions on product clips: use korean-ad, not slam
slam, punch and tiktok-green have no Hangul glyphs and return 400 on Korean copy. Use korean-ad with language ko for Korean product clips; $0.20 a job.
Written by Sume