YouTube B-roll batch: ten 8-second AI clips on Sume, cost and waves

Ten 8-second 16:9 B-roll clips cost $8.00 on Omni 1080p and $10 on Wan 720p. How many waves a Free, Pro or Startup plan needs to render them.

5 min readSume
All posts

Ten 8-second 16:9 B-roll clips for a YouTube video cost $15.00 on Gemini Omni Flash 1.1 at 1080p, $10.00 on Wan 3.0 at 720p, and $5.00 on Wan 3.0 at 480p. Eight seconds is the default length for model: "sume/auto", and it is inside every range in the catalog: Omni 3 to 10 seconds, Wan 2 to 30, H3 5 to 15.

Cost of ten clips

Ten 8-second clips, per-job price rounded up to the cent (catalog rates, read 2026-10-05)
Model and resolutionOne 8 s clipTen clips
gemini-omni-flash-1.1 720p$1.00$10.00
gemini-omni-flash-1.1 1080p$1.50$15.00
minimax-h3-max 768p$0.80$8.00
wan-3.0 480p$0.50$5.00
wan-3.0 720p$1.00$10.00

How long the plan takes to chew it

The admission docs list processing concurrency per plan. Ten accepted jobs fit in a Pro queue (24) and a Startup queue (48), but not a Free workspace (6 accepted, so the seventh submit returns 429 queue_full).

Waves needed for ten jobs (Sume docs, read 2026-10-05)
PlanRunning at onceWaves for 10 jobsFits in accepted capacity?
Free110No: 6 accepted
Pro43Yes: 24 accepted
Startup82Yes: 48 accepted

Request and checks

Send one POST /v1/videos per clip with its own Idempotency-Key, aspect_ratio: "16:9" and the same resolution. Check GET /v1/balance first: each submit reserves its estimated cost, so ten jobs hold the full amount at once on a plan that accepts them. The docs say a video usually takes from 30 seconds to several minutes, so measure your own wave time.

For upload rules, resolution or length limits on the channel side, use YouTube's own help pages; this page covers only what Sume accepts and bills.

Making a B-roll set read as one

Ten independent clips will not match each other by luck. Use one style sentence at the start of every prompt, keep the same resolution and aspect_ratio, and reuse the same model. If you have a reference look, a model that accepts input_references lets you carry it through; check supported_input_references for your model first. Submit a first pair, look at them, then send the other eight. That costs one extra poll cycle and can save a whole batch.

A reasonable order is: price the batch in a script, confirm the balance, send two, review, then send the rest in waves sized by the plan.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

If you are new to the API, start with one clip, one model and the lowest resolution, read the full response once, and only then build a loop around it. Most surprises with video jobs come from fields that were defaulted, not from fields that were set.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume