A 40-clip season on each Sume plan: how many rounds of jobs

Eight episodes of five clips is 40 video jobs. Pro runs 4 at a time, Scale 20. Rounds per plan, plus a submitter that survives queue_full.

5 min readSume
All posts

A season of 8 episodes with 5 clips each is 40 video jobs, and how long it takes to render depends mostly on your plan's concurrency: Pro processes 4 jobs at a time and accepts up to 24, so 40 clips take 10 rounds, while Scale processes 20 at a time and accepts up to 120, so the same season is 2 rounds (read 2026-10-03). The plan changes the waiting, not the price per clip.

Series are the reason this comes up. Metricool reports that TikTok's Next Episode program, run with Amplify, is meant to fund creator-led series (reported, read 2026-10-03). A series is a standing production job, not a one-off, and a team that plans eight episodes at once will hit queue limits the first time it submits everything in a loop.

The limits, per plan

The admission docs separate three numbers. Concurrency is how many generation jobs for a workspace run at once. Accepted capacity is how many may be queued or processing together. Anything beyond accepted capacity is rejected with 429 queue_full, which is a signal to wait, not a failure of your request.

Queueing is by design. A workspace with a concurrency limit of 1 can submit several valid jobs and have them all come back as queued, as long as balance and queue capacity are available. Only one runs at a time. A queued job is not an error.

Rounds needed for 40 video jobs by plan (read 2026-10-03)
PlanConcurrencyAccepted capacityRounds for 40 jobsSubmit all at once?
Free1640No, 6 at a time
Pro42410No, 24 then refill as jobs finish
Startup8485Yes, 40 fit in 48
Scale201202Yes

A submitter that respects the queue

The loop below submits clips one at a time and, on queue_full, waits for retry-after when the header is present and tries the same clip again with the same idempotency key. Reusing the key is what makes the retry safe: it is the same request, so it cannot be billed twice. It stops on 402 insufficient_credits, which no amount of waiting fixes.

Keep clip order stable and keys deterministic, such as s1e03c2 for season 1, episode 3, clip 2, so a restart of the script resumes instead of duplicating.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(key, prompt):
    while True:
        r = requests.post(f"{API}/v1/videos",
            headers={**H, "Idempotency-Key": key},
            json={"model": "seedance-2.5", "prompt": prompt,
                  "duration": 6, "aspect_ratio": "9:16"})
        if r.status_code == 429:
            time.sleep(float(r.headers.get("retry-after", 20)))
            continue
        if r.status_code == 402:
            raise SystemExit("insufficient_credits: stop and check the balance")
        r.raise_for_status()
        return r.json()["request_id"]

jobs = {}
for ep in range(1, 9):
    for clip in range(1, 6):
        key = f"s1e{ep:02d}c{clip}"
        jobs[key] = submit(key, f"Episode {ep}, clip {clip}: see the shot list")
print(len(jobs), "jobs accepted")

What to do before you queue a season

Read your own limits first. Generation submit responses include a generation_limits snapshot with the plan's concurrency and accepted-job limits, and the admission docs recommend GET /v1/balance plus those limits for bulk decisions. Do not treat the wave-size hint in that snapshot as concurrency; the docs call it only a submission-wave hint.

Order the queue by what you need first. Submit episode one's clips before episode eight's, because a queued job waits behind earlier ones. If you are reviewing episodes as they finish, a first-in order gives you something to look at while the rest render.

Cancel what you will not use. Cancellation works only before generation work starts, and it releases the reservation. If you rewrite episode five's script while its clips are still queued, cancel those jobs rather than letting them run. If a job has already started, cancel returns a conflict and you pay for it.

Check the balance before the loop. A paid job reserves its estimated cost when it is accepted, so 40 queued jobs hold 40 reservations at once. The balance has to cover the whole accepted set, not just the one running. A wallet that covers four clips can still stop the thirtieth submit with insufficient_credits.

Poll sparingly. Store each job id, poll status with backoff, and fetch the result only once it reports result_ready or completed. Webhooks are an alternative if you would rather not poll at all.

What Sume does and does not do

Sume admits valid jobs up to your accepted capacity, runs them at your concurrency, and rejects the rest with a typed error you can retry. It does not stretch your concurrency because the queue is long, it does not reorder your jobs, and it does not tell you how long a clip will take; the rounds above are a count of waves, not a time estimate. Read the live job status for timing.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume