1000 Sora prompts at 8s 720p 9:16: the bill and the reserve per wave
A 1,000-clip library of 8-second vertical 720p clips costs $600.00 to $2,420.00 on Sume. Sume reserves the cost at submit, so size your waves.

Re-rendering 1,000 prompts as 8-second 720p 9:16 clips costs $600.00 on minimax-h3 at 768p, $1,000.00 on wan-3.0 or gemini-omni-flash-1.1, $1,520.00 on seedance-2-mini and $2,420.00 on seedance-2-fast. Sume reserves each job's estimated cost when it accepts the request, so you do not need the full total in the balance at the first submit.
The Videos API that Sora clients called was removed on 2026-09-24 with no replacement listed by OpenAI, so a thousand-prompt library is a one-time migration job, not a steady feed. Plan it in waves.
The full-library total by model
These totals multiply the per-clip cents by 1,000. They assume 8 seconds, 720p and 9:16 on every prompt.
| Model | Per clip | 1,000 clips |
|---|---|---|
| seedance-2-mini | $1.52 | $1,520.00 |
| seedance-2-fast | $2.42 | $2,420.00 |
| wan-3.0 | $1.00 | $1,000.00 |
| kling-3, audio off | $1.12 | $1,120.00 |
| kling-3, audio on | $1.68 | $1,680.00 |
| minimax-h3 (768p) | $0.60 | $600.00 |
| gemini-omni-flash-1.1 | $1.00 | $1,000.00 |
Reserve, capture, release
Per the generation-admission docs, Sume reserves the estimated amount at submit time when it accepts the request. A successful completion captures the reserved usage, and failed jobs or a failed queue admission release or refund it. If the balance cannot cover the reservation, the submit fails with 402 insufficient_credits before any provider work starts.
That means a wave of 50 jobs on seedance-2-fast holds 50 x 242 cents = $121.00 at once, while the same wave on wan-3.0 holds $50.00. You can run the whole library with a balance far below the library total, as long as each wave fits and the captured spend is topped up as you go.
Choosing a wave size
Two limits shape a wave. The first is the balance that you are willing to hold in reservations. The second is the queue capacity of the workspace: when it is full, a submit returns 429 queue_full, and the right response is to wait or cancel queued jobs and then retry with the same idempotency key.
A simple loop submits a wave, waits for every job to reach a terminal state, writes the results, and then starts the next wave. Because each submit carries an Idempotency-Key derived from the prompt row, a crash in the middle of a wave can be restarted without paying twice.
- Start with a wave of 20 and raise it only while you see no 429 responses.
- Cancel queued jobs you no longer need; cancel succeeds only before generation starts.
- Keep the polling fallback even if you use
callback_url.
Sources
Related posts
More in Pricing
- 1,000 seconds of Omni Flash: Google about $101, Sume $37.50 to $375
1,000 seconds of Gemini Omni Flash is about $101.36 at Google's 720p token rate and $37.50 to $375 on Sume by resolution. The arithmetic is shown.
- 1,000 drafts: GPT Image 2.5 low at $10 vs Seedream 5.0 Lite at $50
A thousand draft images cost $10 on GPT Image 2.5 at quality low and $50 on Seedream 5.0 Lite on Sume. The cents arithmetic and when the dearer row still wins.
- 1,000 transparent PNGs: gpt-image-2.5 background vs model plus RMBG
A transparent PNG costs $0.0074 to $0.066 from gpt-image-2.5 alone, or the model price plus $0.0225 RMBG for other rows. 1,000-image table on Sume.
- 1080p premium by row: Recast +50%, H3 Max +100%, H3 has no 1080p
Moving up to 1080p adds 50% to H3 Max Recast and 100% to MiniMax H3 Max on Sume, while MiniMax H3 stops at 768p. Per-second and per-10-second figures.
Written by Sume