100 holiday videos in one Sume bulk run: how long with concurrency 16

A bulk queue of 100 items at concurrency 16 drains in 7 waves. Wave arithmetic from the docs' 15 to 30 minute long-form figure, and what else caps the window.

5 min readSume
All posts

How long does a 100-video holiday batch take on Sume? The docs do not promise a number for your Format, so the honest answer is arithmetic on a figure you measure. A queue accepts 1 to 100 items and a concurrency of 1 to 16. With 100 items and a window of 16 the server runs 16 children at once and starts the next the moment a slot frees, which is at most 7 waves (100 divided by 16, rounded up). If a run takes T minutes, a perfectly even queue finishes in about 7 times T.

For scale only: the Format API overview says long-form host video typically finishes in 15 to 30 minutes. That would put 7 waves at roughly 105 to 210 minutes. Short clips will be faster, and any number here is a planning aid, not a guarantee.

Why the real figure can be longer

The window is a ceiling on in-flight children, not a promise that 16 run. The Bulk runs page says children still go through ordinary run admission: wallet, workspace generation concurrency and spend caps. If the workspace allows fewer generations at once than your window, the extra children wait, and the wave math understates the time.

Items also finish unevenly. The queue does not wait for a wave to complete; it refills slots one at a time, so one slow video does not block the next fifteen. That helps the average but not the last item, which sets when the queue becomes completed.

A planning table

The rows below are the wave counts for common holiday batch sizes at the maximum window. Multiply by your measured per-run minutes.

Wave arithmetic at concurrency 16, from the Sume Bulk runs docs read 2026-10-04
ItemsConcurrencyWaves (rounded up)
16161
50164
100167
100813
100425

The tail and the retries

The last few items decide when you are done. Because slots refill one at a time, the last wave is often a partial wave, and the queue is completed only when the slowest remaining child is terminal. Set your alert on the queue's finished_at rather than on an average.

Retries are a second queue. A failed item needs a new run, and the docs advise retrying with a new idempotency key or continuing with previous_run_id. Budget for a smaller second queue made of just the failed rows, and put its time on the same calendar.

Finally, remember the 90-minute run ceiling: each child has an expires_at, 90 minutes from creation or sooner when silent, after which it is finalized as failed. A queue of 100 does not extend that per-run limit.

Compute the estimate from your own sample

Run ten items first, take the longest and the median wall time, and feed both into this snippet. It prints a best and worst case, so you see the spread before promising a date.

import math


def estimate(items, concurrency, minutes):
    waves = math.ceil(items / concurrency)
    return waves, waves * minutes


median, longest = 12, 22   # minutes, measured on ten sample runs
for c in (16, 8, 4):
    w, best = estimate(100, c, median)
    _, worst = estimate(100, c, longest)
    print(f"concurrency {c}: {w} waves, {best} to {worst} min")
assert estimate(100, 16, 1)[0] == 7

Plan the calendar backward

Pick the date the videos must be live, subtract the worst case, subtract time to retry failed items, and start there. A queue is completed when every item is terminal, and failed children need a new run, so reserve at least one more full wave for retries. Poll with backoff and read counts for progress; there is no queue webhook.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume