402 on the third submit: six 5-second Wan 3.0 jobs on a $1.50 balance
Each 5 s Wan 3.0 720p job reserves $0.625 at submit. On a $1.50 balance the third POST /v1/videos fails with 402 insufficient_credits. Arithmetic and a script.

On a $1.50 workspace balance, six 5-second Wan 3.0 jobs at 720p get through two submits and fail the third with 402 insufficient_credits. Each job reserves $0.625 when Sume accepts it, so after two jobs only $0.25 is left. The check happens at submit, before provider work starts, and not when the video finishes, so a batch loop meets the 402 early and not late.
The arithmetic
Sume bills the provider list price times 1.25, and the Sume list price for wan-3.0 at 720p is $0.125 per output second (read 2026-10-08). The docs say Sume reserves the estimated amount at submit and captures it when the job completes, so the reservation is the number to compare with your balance. The queue does not change this: the admission guide says Sume creates a job only when the request is valid, it can reserve the balance, and the workspace still has capacity.
For six jobs the full reservation would be 6 x $0.625 = $3.75, which is $2.25 more than the $1.50 balance. A pre-flight comparison of that total with the balance would have shown the shortfall before any job was submitted.
| Submit | Reserve | Balance before | Balance after | Result |
|---|---|---|---|---|
| 1 | 5 x 0.125 = $0.625 | $1.50 | $0.875 | 202 accepted |
| 2 | $0.625 | $0.875 | $0.25 | 202 accepted |
| 3 | $0.625 | $0.25 | unchanged | 402 insufficient_credits |
| 4 to 6 | $0.625 each | $0.25 | unchanged | not submitted |
A check you can run
The script reproduces the table without calling the API. When you adapt it, take the price and the allowed durations from GET /v1/videos/models and the balance from GET /v1/balance. It assumes that open reservations reduce the spendable balance, which is how the admission docs describe the reserve step. wan-3.0 accepts 2 to 30 seconds, so a 5 second request is inside its range.
import asyncio
PRICE_PER_SECOND = 0.125 # wan-3.0 720p, Sume list price (read 2026-10-08)
async def main() -> None:
balance, seconds = 1.50, 5
reserve = round(PRICE_PER_SECOND * seconds, 4)
for n in range(1, 7):
if balance < reserve:
print(f"job {n}: 402 insufficient_credits (need {reserve}, have {balance})")
return
balance = round(balance - reserve, 4)
print(f"job {n}: accepted, spendable balance now {balance}")
asyncio.run(main())What to do on the 402
The admission guide gives the remedy for 402 insufficient_credits: upgrade the plan or wait for included credit, or submit a less expensive request. Do not retry the same body in a loop, because the balance has not changed. This differs from the 429 queue_full error, which means the workspace has no accepted-job capacity left and clears once a running job finishes. For a 402 the retry advice is about the money, for queue_full it is about time.
Also separate the two checks in your own code. A 402 is decided from money and a request, so you can predict it before you submit by multiplying the per-second price by the duration. A queue_full is decided from the state of the workspace at that moment, so you cannot predict it from the request, and the generation_limits snapshot in submit responses is the place to read it. A batch client should therefore compute the total reservation for the whole batch once, compare it with the balance before the first submit, and then watch generation_limits while it submits.
- Shorter duration: a 4 second job reserves 4 x 0.125 = $0.50, so $1.50 covers three of them.
- Lower resolution: wan-3.0 at 480p is $0.0625 per second, so 5 seconds reserves $0.3125 and $1.50 covers four jobs.
- Stop at the first 402 and store the remaining prompts. The jobs already accepted keep running and complete normally.
Sources
Related posts
More in Developers
- 47 voiceover lines, Timeline's 20-part limit: three concats, 3 cents
Timeline audio joins up to 20 parts per job at $0.01. 47 lines need 20 + 20 + 7 = three concats ($0.03). 47 short lines of TTS also bill the 1-cent floor each.
- 48 or 50 fps for YouTube: Sume output.fps accepts 24, 25, 30, 60
YouTube lists 24, 25, 30, 48, 50 and 60 fps as common rates. Sume Timeline's output.fps takes 24, 25, 30 or 60, so omit it for 48 or 50 sources. Why.
- 4K 3840x2160 for YouTube: Sume Timeline output stops at 2160 per edge
YouTube lists 35-45 Mbps for 4K. Sume Timeline allows even width and height from 256 to 2160, so 3840x2160 is refused. What you can render instead.
- 60 clips on hosted MCP: three jobs_wait batches of 20, 3 calls total
jobs_wait takes 1 to 20 job_ids. For 60 clips that is three batch waits instead of 60. Add include_results to skip the separate result reads too.
Written by Sume