Wan 3.0 on Alibaba: 300 requests a minute across all regions, planned

Alibaba caps Wan 3.0 at 300 requests per minute across all six regions, so a region switch does not add headroom. How to plan a batch of 30 s jobs.

4 min readSume
All posts

Alibaba's Wan3.0-Video page gives one rate limit: 300 requests per minute, across all regions. That means moving a job from Singapore to Virginia does not buy a second 300. For a batch, the limit sets how fast you can submit.

What the Alibaba page states

Figures are from the Alibaba Cloud Model Studio page, last updated September 28, 2026.

Wan3.0-Video limits, read 2026-10-05
ItemValue
Rate limit300 requests per minute, shared across all regions
RegionsBeijing, Singapore, Tokyo, Frankfurt, Virginia, Hong Kong
Max clip length30 seconds
Resolutions480P, 720P, 1080P
Price per second at 720P$0.082513 (five regions) or $0.1 (Singapore)

Planning a batch

The page does not say whether a status poll counts as a request, so budget as if it does. This is planning arithmetic, not a vendor figure.

Batch planning against 300 requests per minute, read 2026-10-05
PlanArithmeticResult
Submit 1,000 clips at the ceiling1,000 / 300 per minute3.33 minutes of submit time
Poll every 15 s if polls count4 polls per minute per job75 jobs in flight at 300 per minute
Poll every 30 s if polls count2 polls per minute per job150 jobs in flight at 300 per minute
Cost of 1,000 clips of 30 s at 720P1,000 x 30 x $0.082513$2,475.39 in the five regions

The same plan on Sume

Sume's video API is asynchronous: POST /v1/videos returns 202 with an id and a polling URL, then you poll GET /v1/videos/{id}, or send a callback_url and let Sume call you when the job ends. The Sume error table has a 429 rate_limited response. The Sume docs do not state a requests-per-minute number, so this post does not quote one. Handle a 429 with backoff, as below.

A webhook removes the poll traffic from your count. Sume signs the callback body and sends x-sume-webhook-signature, and the webhook payload is the standard Sume job envelope.

import asyncio, os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

async def submit(prompt):
    body = {"model": "wan-3.0", "prompt": prompt, "duration": 30, "resolution": "720p"}
    for attempt in range(5):
        r = await asyncio.to_thread(
            requests.post, "https://api.sume.com/v1/videos", headers=H, json=body
        )
        if r.status_code == 429:
            await asyncio.sleep(2 ** attempt)
            continue
        r.raise_for_status()
        return r.json()["id"]
    raise RuntimeError("still rate limited after 5 tries")

async def main():
    prompts = [f"Aerial shot of a coastal road, scene {i}" for i in range(3)]
    print(await asyncio.gather(*(submit(p) for p in prompts)))

asyncio.run(main())

Checklist

Keep a queue in front of the API, limit in-flight jobs, and treat 429 as normal flow control.

  • Use a callback or a slow poll interval; the Sume docs suggest about 30 seconds.
  • Cap submissions per minute below the limit you know about.
  • Send an Idempotency-Key on Sume so a retry returns the original job and does not make a second one.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume