Wan 2.2 prompt extension: a DashScope key or a local Qwen model
Wan 2.2's README offers prompt extension via Alibaba DashScope (DASH_API_KEY) or a local Qwen model. What each needs, and what Sume's request takes.

Running Wan 2.2 weights does not mean running without any hosted service. The README's optional prompt extension rewrites your prompt with a language model, and it offers two ways: Alibaba's DashScope API, which needs a DASH_API_KEY, or a local Qwen model from Hugging Face. Leave it off and Wan 2.2 uses your prompt as written. Sume's video request has a prompt field and no prompt-extension field in the documented parameters. Details read from the Wan2.2 README on 2026-10-08.
The two extension routes
Both routes are switched on with --use_prompt_extend and chosen with --prompt_extend_method.
| Route | What you need | Models the README names |
|---|---|---|
| dashscope | A DashScope API key in DASH_API_KEY; international site users also set DASH_API_URL | qwen-plus for text-to-video, qwen-vl-max for image-to-video |
| local_qwen | A Qwen model on disk or from Hugging Face; more GPU memory for a larger one | Qwen2.5-14B, 7B or 3B Instruct for text-to-video; Qwen2.5-VL-7B or 3B for image-to-video |
What it costs you in effort
The DashScope route puts a second account and a second bill in your pipeline. The local route puts a second model in GPU memory: the README says larger models give better extension but need more GPU memory, and you can point --prompt_extend_model at a local path. On a 24 GB card running the 5B model, that competes with the video model for memory.
The README's example commands use eight GPUs with torchrun --nproc_per_node=8 for the A14B models, so the extension step is one more thing to schedule on that node.
Where the prompt step fits
The README places prompt extension before generation: the extension model reads your short prompt and returns a longer one, which the video model then uses. With DashScope the call leaves your machine; with local Qwen it stays on your hardware. That is a data-handling choice as well as a cost choice, and the README does not state a retention policy for DashScope, so read Alibaba's terms for that service.
Without extension, write the long prompt yourself: subject, motion, camera, lighting.
The hosted comparison
A Sume video request takes a prompt and optional fields (duration, resolution, aspect_ratio, frame_images, input_references). A minimal request for the hosted Wan model is:
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wan-prompt-001" \
-d '{
"model": "wan-3.0",
"prompt": "A paper boat drifts down a rain gutter, low camera",
"resolution": "720p",
"duration": 5
}'Sume lists wan-3.0, not Wan 2.2, and it accepts 2 to 30 seconds. At 720p a five-second clip is $0.625 (list $0.10 per second x 1.25). If you want better prompts, write them yourself or use another tool; Sume's video docs advise detailed prompts covering motion, camera, lighting and composition, and do not describe an automatic rewrite step.
Which to pick
Use DashScope if you accept a second vendor and want the larger hosted model without spending GPU memory on it. Use local Qwen if prompts must stay on your hardware or you already hold the Hugging Face weights. Skip extension if your prompts are already detailed, since it is optional in the README and each route adds a failure point: an invalid key, a missing model path, or an out-of-memory error from loading two models at once.
- DashScope: needs
DASH_API_KEY; setDASH_API_URLif you use the international site. - Local Qwen: pick a model size that fits next to the video model in memory.
- Neither: leave
--use_prompt_extendoff.
Sources
Related posts
More in Models
- Wan 3.0 leads text-to-video at 1,156 Elo: $12 listed, $15 on Sume
Wan 3.0 is first on the AA-Video-T2V v2.0 board at 1,156 Elo and $12.00 a minute. On Sume, wan-3.0 at 1080p bills $0.25 a second, which is $15.00 a minute.
- Wan 3.0 reference limits: 10 images, 5 videos and 5 audio tracks
Sume lists Wan 3.0 references as up to 10 images, 5 videos (15 s total) and 5 audio tracks (15 s total). Sume docs do not claim 50 references.
- What a 13-point Elo gap means: Wan 3.0 vs Seedance 2.5 win chance
Wan 3.0 leads Seedance 2.5 by 13 Elo on the AA text-to-video board. On the standard Elo scale that is a 51.9% win chance per vote. Arithmetic for other gaps.
- When Seedance 2 Mini or Fast is enough, and when to pay for 2.5
A decision guide for Seedance tiers on Sume: what Mini, Fast and 2.5 cost for the same clip, what only 2.5 can do, and which jobs belong on which tier.
Written by Sume