Video batch budget guard in Python: stop when usage.cost hits the cap
Run a batch of Sume video jobs one at a time, add up each finished job's usage.cost, and stop before the next submit would cross a dollar cap.

Short answer
Submit one video job, poll until it reaches a terminal state, add usage.cost to a running total, and refuse to submit the next job if the total plus that job's estimate would pass your cap. Sume reserves provider list times 1.25 at submit and usage.cost is the billable amount, per the video docs. Because the guard checks before each submit, an overrun is bounded by one job's estimate.
The guard
This runs Wan 3.0 drafts at 480p and 5 seconds, which is $0.0625 a second, so $0.3125 per clip. It polls every 10 seconds. Set the cap and the prompt list for your batch.
import asyncio, json, os, urllib.request
BASE = "https://api.sume.com/v1/videos"
HDR = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
CAP, EST = 2.00, 0.3125
def call(url, body=None):
data = json.dumps(body).encode() if body else None
with urllib.request.urlopen(urllib.request.Request(url, data, HDR)) as r:
return json.load(r)
async def one(prompt):
body = {"model": "wan-3.0", "prompt": prompt, "duration": 5,
"resolution": "480p"}
job = await asyncio.to_thread(call, BASE, body)
while job["status"] not in ("completed", "failed", "cancelled"):
await asyncio.sleep(10)
job = await asyncio.to_thread(call, job["polling_url"])
return job
async def main():
total = 0.0
for prompt in ["A kite over dunes", "A neon ramen stall"]:
if total + EST > CAP:
break
job = await one(prompt)
total += (job.get("usage") or {}).get("cost", 0)
print(job["status"], round(total, 4))
asyncio.run(main())Why one at a time
Running jobs in parallel is faster but each one reserves its estimate at submit, so a cap check made before submitting eight at once cannot see the other seven. Sequential submits keep the total honest. If you need speed, split the cap: give each worker its own slice of the budget and apply the same guard inside it.
What counts toward the total
| Field or event | Meaning | Use in the guard |
|---|---|---|
| Reserve at submit | Provider list x 1.25 held when you submit | Use as the per-job estimate |
| usage.cost | Sume billable amount on the finished job | Add to the running total |
| failed job | Reserve is refunded | usage.cost may be absent; the guard adds 0 |
| Idempotency-Key replay | Returns the original job | Does not add a second charge |
Improving it
Add an Idempotency-Key per prompt so a retry after a timeout returns the original job instead of a second one. Log the job id with the cost so you can reconcile against your balance. Swap EST for the estimate of whichever model you use; for a Seedance 2.5 clip it is the token arithmetic, not a flat per-second rate.
Testing the guard
Set the cap to $0.40 with a $0.3125 estimate and a three-prompt list, and run it: the first clip should run and the second should be refused because $0.3125 plus $0.3125 passes $0.40. If both run, your total is not being updated. Then run it with a model id that does not exist and confirm the error is caught early. Cheap tests like this are worth doing before a batch that costs real money.
Caveats
The cap counts finished jobs only; an in-flight job's reserve is held but not yet in total, which the EST margin covers. If the process dies mid-batch, the job continues on the server, so keep job ids to reconcile.
Related posts
More in Developers
- v1/videos/models `created` is a catalog date, not a release date
Every model on Sume's /v1/videos/models shows created 1767225600, which is 2026-01-01. It is not when Gemini Omni 1.1 Flash or MiniMax H3 launched.
- Video filter stops at 300 seconds: which short-video lengths pass?
Sume's video filter refuses sources over 300 seconds. A 3-minute Short passes; a 10-minute TikTok ad source does not and needs a trim first.
- Video Router or /v1/videos for ported Sora code?
Both routes create the same Sume video jobs with the same model ids. /v1/videos is the OpenRouter-style wire; /v1/video-router/generate is the older flat one.
- Voice agent hand-off: ask for a clip, get an async Sume job
A live voice agent should not wait on a video render. Hand the request to an async Sume job, speak the job id back, and deliver the clip by poll or webhook.
Written by Sume