300 holiday UGC clips on one plan: pace around queue_full 429s

A Pro workspace holds 4 processing and 20 queued jobs. Submit 300 avatar clips, treat queued as normal, and retry queue_full with the same idempotency key.

5 min readSume
All posts

You can submit far more clips than your workspace processes at once, but only up to its accepted-job capacity: on the Pro plan that is 4 processing plus 20 queued, so 24 paid generation jobs at a time. The 25th returns 429 queue_full, and the documented response is to wait for jobs to finish, then retry with the same idempotency key. A 300-clip UGC batch is therefore a paced submit loop, not one burst.

Treat queued as normal. The docs say concurrency is a dispatch limit, not a submit limit.

What are the limits for each plan?

Generation concurrency is plan-only; prepaid top-ups do not raise it. Queue capacity defaults to max(3, concurrency_limit x 5), and accepted capacity is the sum of the two. Org workspaces have a floor of 10 and Enterprise uses admin overrides, so read the effective generation_limits.concurrency_limit in your responses or the dashboard Concurrency tab rather than trusting this table.

For 300 clips, accepted capacity divided into the batch gives the minimum number of waves. On Pro that is 300 / 24, about 13 waves. How long each wave takes depends on the clips; the docs give no per-job duration, so we do not estimate one.

Default generation admission by plan (Sume docs, read 2026-10-02)
PlanProcessingQueue capacityAccepted at onceWaves for 300 clips
Free15650
Pro4202413
Startup840487
Scale201001203

Which errors mean wait, and which mean stop?

Four codes show up in a peak batch and they need different handling. queue_full and rate_limited are both 429, but the first means capacity and the second means request volume. insufficient_credits is a 402 returned before provider work starts, because Sume reserves the estimated cost at submit; retrying will not fix it.

A failed admission releases or refunds its reservation where applicable. Do not resubmit a paid request just because your own process timed out; look up the job first. That is why every submit below carries a stable Idempotency-Key per clip.

  • queue_full (429): stop adding work, poll or cancel, retry with the same key after capacity opens.
  • rate_limited (429): back off using retry-after when present.
  • insufficient_credits (402): the reservation could not be funded; lower the request cost or fund the workspace per the docs, and do not loop.
  • idempotency_conflict (409): the key was reused for a different payload.

What does the paced submit loop look like?

This script submits 300 avatar clips with a per-clip key and sleeps on queue_full. It uses only the standard library and runs as is with SUME_API_KEY set; the avatar handle is the one used in the Sume docs examples, so replace it with your own ready avatar. It does not poll results, which you should do through job status or webhooks.

Because every clip has its own key, re-running the whole script after a crash is safe: a replayed key returns the original job instead of charging again.

import json, os, time, urllib.error, urllib.request

KEY = os.environ["SUME_API_KEY"]
URL = "https://api.sume.com/v1/avatar-1.0/talking-video"

def submit(n):
    body = {"avatar_handle": "sume_clawra", "aspect_ratio": "9:16",
            "script": f"Gift idea number {n}: this one arrives wrapped, ships fast, and fits any budget this holiday season."}
    req = urllib.request.Request(URL, json.dumps(body).encode(), {
        "Authorization": f"Bearer {KEY}",
        "Content-Type": "application/json",
        "Idempotency-Key": f"peak-clip-{n}"})
    while True:
        try:
            with urllib.request.urlopen(req) as r:
                return json.load(r)
        except urllib.error.HTTPError as e:
            err = json.load(e).get("error", {})
            if e.code == 429 and err.get("code") == "queue_full":
                time.sleep(int(e.headers.get("retry-after", 30)))
                continue
            raise

for n in range(1, 301):
    submit(n)

What should you not use `wave_size_hint` for?

Submit responses can include a generation_limits snapshot with a wave_size_hint, defined as max(1, floor(queue_capacity_remaining x 0.75)). On an idle Pro workspace that is 18. The docs are explicit that it is a submission-wave hint only: not a concurrency limit and not a processing width, so never size in-flight work with it.

If a sale moment changes and you no longer need the tail of the batch, cancel jobs that are still queued. Once generation has started, cancel returns 409 job_generation_already_started and the job finishes or fails normally.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume