Where to use Wan 3.0: wan.video, QwenCloud, or the wan-3.0 API on Sume
Which door to Wan 3.0 fits: Alibaba's channels named in its README, or the wan-3.0 id on Sume with an API key, MCP and Timeline for multi-shot work.

There are two kinds of doors to Wan 3.0. Alibaba's own channels, which its README (read 2026-10-05) names as wan.video and QwenCloud, and third-party APIs such as the wan-3.0 model id on Sume. Choose Alibaba's channels if you want Alibaba's own features exactly as Alibaba ships them, including document and webpage references. Choose Sume if you want one API key and one job format for Wan 3.0 next to other video models, plus Sume's media tools for what comes after generation. The model is not open weights: the README and the launch coverage describe a hosted release, so there is no file to download and run yourself.
What the sources say about Wan 3.0
The TechNode report on the launch (read 2026-10-05) is dated 24 August 2026 and describes 30 second generation and document input. The README lists the features: native 30 seconds, up to 20 reference assets (documents, webpages, text, images), instruction- and reference-based video editing, automatic scene splitting, pixel-level box-selection editing, up to 12 sequential images with a cohesive style, and text rendering in 12 languages.
What Sume exposes
Sume's catalog id is wan-3.0. The Video generation docs give its length range as 2 to 30 seconds and say it accepts image, video and audio references. The catalog constraints add: up to 10 images, 5 videos and 5 audio files, and no web_url or file_url field in v1. Video-to-video editing is not listed as a capability for this model, so Alibaba's instruction-based editing and box selection are not request fields on Sume.
| You need | Better fit | Why |
|---|---|---|
| Documents or webpages as direct input | Alibaba's channels | README lists them as reference assets; Sume has no such field |
| Box-selection or instruction-based editing | Alibaba's channels | Not exposed on Sume for wan-3.0 |
| Wan 3.0 beside other video models in one API | Sume | One key, one job format, catalog read from GET /v1/videos/models |
| Join shots, trim, caption, inspect the result | Sume | Timeline 1.0, video trim, video captions, video inspect |
| Call it from Claude or Cursor | Sume hosted MCP | generate_video tool, listed in the tools and gates page |
| Run the weights on your own GPUs | Neither today | Closed beta and API; no open weights |
Reading this honestly
This post does not compare prices or quality between the channels. I did not read the wan.video or QwenCloud pages in detail, only what the README says about their existence, so any detail about those products should come from Alibaba. For Sume, price is read at call time from pricing_skus in GET /v1/videos/models, which is the only number to plan with.
- If the feature you want is on the README list but not in this table, ask whether it is a request field or a model behavior.
- If you start on Sume and later need a feature Sume does not expose, nothing stops you from using Alibaba's channel for that one job.
- Keep your prompts in files, not in a UI, so they move between doors.
A quick test
Run the same prompt, 5 seconds at 480p, through each door you are considering. Compare the clips, the fields you could set and the time to a usable file. That small experiment is cheaper than a long evaluation, and it shows whether the limits above matter for your work. For Sume the request is one POST /v1/video-router/generate with model: "wan-3.0", a prompt, duration and resolution, and an Idempotency-Key header.
Sources
Related posts
More in Comparisons
- Which MAI-Voice-2.1 model for ad voiceover: standard or Flash?
Microsoft positions MAI-Voice-2.1 for voice-over and audiobooks, Flash for live agents. For rendered ads pick standard; Sume is the async file path.
- Text change, region change or background swap: which Sume image route
Pick the Sume image edit route by the kind of change: Ideogram 4.5 for words, ChatGPT Image 2.5 with mask_url for a region, a reference edit for backgrounds.
- Which platforms auto-detect AI media: YouTube, Meta, Pinterest, TikTok
YouTube, Meta and Pinterest say they can label AI media from metadata or detection; TikTok auto-labels its own AI effects. None replaces your own disclosure.
- Who owns a Suno song? What 'Output owned by Suno' means for ads
Suno's terms assign Pro and Premier users the Output Suno owns, but make no copyright warranty. What to file with an ad, and how a Sume music job record helps.
Written by Sume