Wan 3.0 API rate limit: 300 RPM on Model Studio vs Sume jobs

Alibaba lists 300 requests per minute for Wan 3.0 on Model Studio. Here is what that means for a batch, and how Sume submits Wan 3.0 as async jobs.

4 min readSume
All posts

Alibaba's Wan 3.0 page lists a limit of 300 requests per minute (RPM) for the model on Model Studio. That is a per-account request ceiling on Alibaba's side, so a batch that submits faster than 5 jobs a second will see throttling there. On Sume you call wan-3.0 through the async video API, and this post does not claim a Sume RPM figure because the docs I read do not publish one for this model.

What Alibaba publishes

The Wan 3.0 page accepts text, image, video and audio input at 480P, 720P and 1080P for up to 30 seconds. It lists six regions: Beijing, Singapore, Tokyo, Frankfurt, Virginia and Hong Kong. Prices differ by region, for example 720P at $0.082513 per second in Beijing, Tokyo, Frankfurt, Virginia and Hong Kong against $0.10 in Singapore.

Why 300 RPM rarely binds on a video job

A 30-second 1080P clip is one request that runs for minutes, so 300 RPM is only a limit on how fast you can start work, not on how many jobs can run. The practical ceiling for most teams is how many generations you can pay for at once, which is a budget question first.

How Sume handles the same batch

Sume's POST /v1/videos returns immediately with a job id, a polling_url and status pending. You poll until the status is completed or failed, or pass an HTTPS callback_url to be notified. The model id is wan-3.0, accepting 2 to 30 seconds at 480p, 720p and 1080p.

Wan 3.0 request handling, Alibaba and Sume (read 2026-10-07)
ItemAlibaba Model StudioSume
Model idWan3.0-Video (see page)wan-3.0
Request limit published300 RPMNot published in the docs read for this post
Clip lengthUp to 30 s2 to 30 s
Result deliveryPer Alibaba docsPoll polling_url or HTTPS callback_url
Rate basisRegion price listVendor list x 1.25

What to do

For a large batch, queue on your side, submit at a steady pace and store the job id from each response. If you need Alibaba's regional prices or its Singapore versus Virginia split, call Alibaba directly.

A pacing plan for a 600-job batch

Say you need 600 clips. At 300 RPM, Alibaba's ceiling lets you start all of them in two minutes, so the limit is not what slows you down; what slows you down is the generation time and your spend. A queue that submits 20 jobs, waits for completions, then submits the next 20 keeps the spend visible and makes failures cheap to find.

  • Store the job id with your own record id the moment the submit call returns.
  • Poll at a modest interval; the Sume docs example waits 30 seconds between polls.
  • Treat failed as a state to log and retry once, not to loop on forever.
  • Set an HTTPS callback_url if you would rather be called than poll.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume