Queue a 100-episode Shorts season in one Sume call?

Yes: one bulk-runs call queues up to 100 Format runs, 1 to 16 at a time. Here is what it returns, what it skips, and how to poll a season.

4 min readSume
All posts

Yes. A bulk request to a Sume Format queues 1 to 100 Format runs in one call, and a concurrency window of 1 to 16 decides how many run at the same time. That fits one Shorts season, or two seasons of 50 episodes, per request.

YouTube's Help page on Shorts series describes a series as a Shorts-only playlist with seasons and episodes, with each episode at most 180 seconds (read 2026-10-07). The queue is how you turn that structure into a batch: one item per episode, in order.

What one bulk call accepts

The route is POST /v1/formats/{handle}/{slug}/bulk-runs (the opaque skl_ path is a twin). The body has two keys. concurrency is an integer from 1 to 16. items is an array of 1 to 100 entries, and each entry is the same body as a single Format run. Sume validates every item before it creates the queue, so one bad item returns 400 invalid_request with details.index and dispatches nothing.

Each item must name at least one of instruction, input, previous_run_id or attachments. For a season, the natural shape is one item per episode with the episode number, the logline and the cliffhanger in the instruction.

Bulk queue limits from the Sume docs (read 2026-10-07)
ControlValueWhat it means for a season
Items per queue1 to 100One call holds a 100-episode season
Concurrency window1 to 16How many episodes render at once
Success status202The first window of items is already in flight
Queue-level webhookNonePoll the queue; set webhooks on each item
Cancel a queueNot offeredCancel children one by one with the run cancel route

Poll the season, not each episode

The queue has no webhook of its own. You poll GET /v1/format-run-queues/{queue_id} for progress, and each child run has its own receipt at GET /v1/format-runs/{run_id}. A per-item communication.webhook_url is the way to get a push when one episode finishes.

Two traps are worth writing down before the first launch. completed on the queue means that every item is terminal, not that every item succeeded, so branch on counts.failed. And mint a fresh Idempotency-Key for each batch: replaying a spent key returns 202 with the old queue, which is useful for a retry and wrong for season two.

  • Send Idempotency-Key on the create so a network retry does not start a second season.
  • Keep concurrency at or below your plan's generation concurrency; the extra runs wait as queued work anyway.
  • Put the season and episode number in each item so the receipts are easy to match to a playlist order.

Where the limits come from

Workspace generation concurrency still applies to the children, so a window of 16 on a small plan does not mean 16 simultaneous renders. See the Free-plan waves post for how accepted capacity caps a batch, and the cost post for what ten episodes cost. Service-account keys cannot create Format runs or bulk queues; they fail with 403 insufficient_scope.

Check the Bulk runs page for the current schema before you script against it, because the live OpenAPI document is the source of truth.

A season plan that survives a bad night

Write the season as a spreadsheet before you write any request: one row per episode, with the number, the logline, the cliffhanger, the model, the resolution and the spend cap. The queue body is then a loop over the rows, and every receipt matches a row. When episode 37 fails, you know which line of the sheet to fix.

Plan for partial failure. A queue of 100 will rarely finish with every child successful, so budget a second pass. Collect the failed indexes from the queue, build a new items array that holds only those episodes, and submit it with a fresh Idempotency-Key. Each retried item can carry previous_run_id when you want the agent to keep the context of the first try, as the scene redo post explains.

Finally, pick the window with intent. A window of 2 is slow and gentle on the workspace. A window of 16 finishes fast, but the children still share the workspace's accepted generation capacity, and a full workspace answers new submissions with 429 queue_full (see the retry post). Keeping the window near the number of generation slots on your plan avoids waiting on that limit.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume