1000 Sora prompts at 8s 720p 9:16: the bill and the reserve per wave

A 1,000-clip library of 8-second vertical 720p clips costs $600.00 to $2,420.00 on Sume. Sume reserves the cost at submit, so size your waves.

5 min readSume
All posts

Re-rendering 1,000 prompts as 8-second 720p 9:16 clips costs $600.00 on minimax-h3 at 768p, $1,000.00 on wan-3.0 or gemini-omni-flash-1.1, $1,520.00 on seedance-2-mini and $2,420.00 on seedance-2-fast. Sume reserves each job's estimated cost when it accepts the request, so you do not need the full total in the balance at the first submit.

The Videos API that Sora clients called was removed on 2026-09-24 with no replacement listed by OpenAI, so a thousand-prompt library is a one-time migration job, not a steady feed. Plan it in waves.

The full-library total by model

These totals multiply the per-clip cents by 1,000. They assume 8 seconds, 720p and 9:16 on every prompt.

1,000 clips at 8 s 720p 9:16, billable on Sume (read 2026-10-08)
ModelPer clip1,000 clips
seedance-2-mini$1.52$1,520.00
seedance-2-fast$2.42$2,420.00
wan-3.0$1.00$1,000.00
kling-3, audio off$1.12$1,120.00
kling-3, audio on$1.68$1,680.00
minimax-h3 (768p)$0.60$600.00
gemini-omni-flash-1.1$1.00$1,000.00

Reserve, capture, release

Per the generation-admission docs, Sume reserves the estimated amount at submit time when it accepts the request. A successful completion captures the reserved usage, and failed jobs or a failed queue admission release or refund it. If the balance cannot cover the reservation, the submit fails with 402 insufficient_credits before any provider work starts.

That means a wave of 50 jobs on seedance-2-fast holds 50 x 242 cents = $121.00 at once, while the same wave on wan-3.0 holds $50.00. You can run the whole library with a balance far below the library total, as long as each wave fits and the captured spend is topped up as you go.

Choosing a wave size

Two limits shape a wave. The first is the balance that you are willing to hold in reservations. The second is the queue capacity of the workspace: when it is full, a submit returns 429 queue_full, and the right response is to wait or cancel queued jobs and then retry with the same idempotency key.

A simple loop submits a wave, waits for every job to reach a terminal state, writes the results, and then starts the next wave. Because each submit carries an Idempotency-Key derived from the prompt row, a crash in the middle of a wave can be restarted without paying twice.

  • Start with a wave of 20 and raise it only while you see no 429 responses.
  • Cancel queued jobs you no longer need; cancel succeeds only before generation starts.
  • Keep the polling fallback even if you use callback_url.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume