Wan 3.0 API rate limit: 300 RPM on Model Studio vs Sume jobs
Alibaba lists 300 requests per minute for Wan 3.0 on Model Studio. Here is what that means for a batch, and how Sume submits Wan 3.0 as async jobs.

Alibaba's Wan 3.0 page lists a limit of 300 requests per minute (RPM) for the model on Model Studio. That is a per-account request ceiling on Alibaba's side, so a batch that submits faster than 5 jobs a second will see throttling there. On Sume you call wan-3.0 through the async video API, and this post does not claim a Sume RPM figure because the docs I read do not publish one for this model.
What Alibaba publishes
The Wan 3.0 page accepts text, image, video and audio input at 480P, 720P and 1080P for up to 30 seconds. It lists six regions: Beijing, Singapore, Tokyo, Frankfurt, Virginia and Hong Kong. Prices differ by region, for example 720P at $0.082513 per second in Beijing, Tokyo, Frankfurt, Virginia and Hong Kong against $0.10 in Singapore.
Why 300 RPM rarely binds on a video job
A 30-second 1080P clip is one request that runs for minutes, so 300 RPM is only a limit on how fast you can start work, not on how many jobs can run. The practical ceiling for most teams is how many generations you can pay for at once, which is a budget question first.
How Sume handles the same batch
Sume's POST /v1/videos returns immediately with a job id, a polling_url and status pending. You poll until the status is completed or failed, or pass an HTTPS callback_url to be notified. The model id is wan-3.0, accepting 2 to 30 seconds at 480p, 720p and 1080p.
| Item | Alibaba Model Studio | Sume |
|---|---|---|
| Model id | Wan3.0-Video (see page) | wan-3.0 |
| Request limit published | 300 RPM | Not published in the docs read for this post |
| Clip length | Up to 30 s | 2 to 30 s |
| Result delivery | Per Alibaba docs | Poll polling_url or HTTPS callback_url |
| Rate basis | Region price list | Vendor list x 1.25 |
What to do
For a large batch, queue on your side, submit at a steady pace and store the job id from each response. If you need Alibaba's regional prices or its Singapore versus Virginia split, call Alibaba directly.
A pacing plan for a 600-job batch
Say you need 600 clips. At 300 RPM, Alibaba's ceiling lets you start all of them in two minutes, so the limit is not what slows you down; what slows you down is the generation time and your spend. A queue that submits 20 jobs, waits for completions, then submits the next 20 keeps the spend visible and makes failures cheap to find.
- Store the job id with your own record id the moment the submit call returns.
- Poll at a modest interval; the Sume docs example waits 30 seconds between polls.
- Treat
failedas a state to log and retry once, not to loop on forever. - Set an HTTPS
callback_urlif you would rather be called than poll.
Sources
Related posts
More in Developers
- Hackathon app on the Sume Free plan: 1 seat, 6 accepted jobs
A weekend demo on Free can have one job processing and five queued. How to design the UI, the retries and the demo script around that, with the real limits.
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
- What is a partial transcript in streaming speech to text?
A partial is a provisional transcript a streaming model revises as audio arrives. Why subtitles for a finished clip only need final text and word times.
- What to show a viewer while an avatar video job is queued
Avatar jobs on Sume are async: queued, processing, then completed, failed or canceled. A status-to-UI map for waiting screens, with polling rules.
Written by Sume