One prompt, two models: asyncio.gather Wan and Seedance for $21.08
Submit one 30 s prompt to wan-3.0 ($3.75) and seedance-2.5 ($17.334) with httpx and asyncio.run. Use one Idempotency-Key per request; a shared key conflicts.

Submitting one 30 s prompt to both wan-3.0 and seedance-2.5 at 720p reserves $3.75 + $17.334 = $21.084, and asyncio.gather sends both requests at once. Give each request its own Idempotency-Key: reusing one key with a different body returns 409 idempotency_conflict.
What the comparison costs
Both models accept 30 s at 720p. Seedance is token-priced, so its total depends on resolution, length and aspect, and the figure below is for 9:16. Wan 3.0 is a flat per-second rate.
Run it once per prompt to choose a model, not on every request; the Seedance side is 4.6 times the Wan side.
| Model | Rate per second | 30 s total | Share of $21.084 |
|---|---|---|---|
| wan-3.0 | $0.125 | $3.75 | 17.8% |
| seedance-2.5 (9:16) | $0.5778 | $17.334 | 82.2% |
| Both | $21.084 | 100% |
The script
Python has no top-level await in a script, so main() runs under asyncio.run. The result lists the job id and polling_url for each model; poll them as in the quick start.
import asyncio, os, uuid
import httpx
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A street food cart at dusk, steam rising, handheld camera"
BODIES = [
{"model": "wan-3.0", "prompt": PROMPT, "duration": 30, "resolution": "720p"},
{"model": "seedance-2.5", "prompt": PROMPT, "duration": 30,
"resolution": "720p", "aspect_ratio": "9:16"},
]
async def submit(client, body):
r = await client.post(
"https://api.sume.com/v1/videos", json=body,
headers={**H, "Idempotency-Key": str(uuid.uuid4())},
)
return body["model"], r.status_code, r.json()
async def main():
async with httpx.AsyncClient(timeout=30) as client:
for model, code, data in await asyncio.gather(*(submit(client, b) for b in BODIES)):
print(model, code, data.get("id"), data.get("polling_url"))
asyncio.run(main())Gotchas
Two paid jobs run at once, so on the Free plan (1 processing, 5 queued) a second pair of submits can wait in the queue, and a seventh paid job gets 429 queue_full. If one side returns 402, the other may already be reserved; check each status code, not just the first.
Keep the generated key with the model name in your own log, so a retry after a timeout reuses it and does not pay twice.
Sources
Related posts
More in Developers
- One webhook route for OpenRouter video and Sume job events
Normalize OpenRouter video.generation.* and Sume job.* webhook bodies to one outcome type, and answer 204 to events you do not know. Runnable Bun/Node code.
- OpenAI's three tiers vs Sume plan concurrency of 1, 4, 8 and 20
OpenAI cut API usage tiers from five to three on Oct 6. Sume sets concurrency by plan, not spend. The plan numbers and the queue math, side by side.
- OpenRouter video client on Sume: cancelled is terminal, callback_url
Moving an OpenRouter video client to Sume: add cancelled to your terminal statuses, send callback_url on each request, and keep unknown statuses non-terminal.
- Org workspace concurrency floor of 10: the queue and wave that follow
Sume gives org workspaces a processing floor of 10. With the default queue formula that is 50 queued, 60 accepted, and a wave hint of 45. Read your own fields.
Written by Sume