Wan 2.2 prompt extension: a DashScope key or a local Qwen model

Wan 2.2's README offers prompt extension via Alibaba DashScope (DASH_API_KEY) or a local Qwen model. What each needs, and what Sume's request takes.

5 min readSume
All posts

Running Wan 2.2 weights does not mean running without any hosted service. The README's optional prompt extension rewrites your prompt with a language model, and it offers two ways: Alibaba's DashScope API, which needs a DASH_API_KEY, or a local Qwen model from Hugging Face. Leave it off and Wan 2.2 uses your prompt as written. Sume's video request has a prompt field and no prompt-extension field in the documented parameters. Details read from the Wan2.2 README on 2026-10-08.

The two extension routes

Both routes are switched on with --use_prompt_extend and chosen with --prompt_extend_method.

Wan 2.2 prompt extension options, README read 2026-10-08
RouteWhat you needModels the README names
dashscopeA DashScope API key in DASH_API_KEY; international site users also set DASH_API_URLqwen-plus for text-to-video, qwen-vl-max for image-to-video
local_qwenA Qwen model on disk or from Hugging Face; more GPU memory for a larger oneQwen2.5-14B, 7B or 3B Instruct for text-to-video; Qwen2.5-VL-7B or 3B for image-to-video

What it costs you in effort

The DashScope route puts a second account and a second bill in your pipeline. The local route puts a second model in GPU memory: the README says larger models give better extension but need more GPU memory, and you can point --prompt_extend_model at a local path. On a 24 GB card running the 5B model, that competes with the video model for memory.

The README's example commands use eight GPUs with torchrun --nproc_per_node=8 for the A14B models, so the extension step is one more thing to schedule on that node.

Where the prompt step fits

The README places prompt extension before generation: the extension model reads your short prompt and returns a longer one, which the video model then uses. With DashScope the call leaves your machine; with local Qwen it stays on your hardware. That is a data-handling choice as well as a cost choice, and the README does not state a retention policy for DashScope, so read Alibaba's terms for that service.

Without extension, write the long prompt yourself: subject, motion, camera, lighting.

The hosted comparison

A Sume video request takes a prompt and optional fields (duration, resolution, aspect_ratio, frame_images, input_references). A minimal request for the hosted Wan model is:

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wan-prompt-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A paper boat drifts down a rain gutter, low camera",
    "resolution": "720p",
    "duration": 5
  }'

Sume lists wan-3.0, not Wan 2.2, and it accepts 2 to 30 seconds. At 720p a five-second clip is $0.625 (list $0.10 per second x 1.25). If you want better prompts, write them yourself or use another tool; Sume's video docs advise detailed prompts covering motion, camera, lighting and composition, and do not describe an automatic rewrite step.

Which to pick

Use DashScope if you accept a second vendor and want the larger hosted model without spending GPU memory on it. Use local Qwen if prompts must stay on your hardware or you already hold the Hugging Face weights. Skip extension if your prompts are already detailed, since it is optional in the README and each route adds a failure point: an invalid key, a missing model path, or an out-of-memory error from loading two models at once.

  • DashScope: needs DASH_API_KEY; set DASH_API_URL if you use the international site.
  • Local Qwen: pick a model size that fits next to the video model in memory.
  • Neither: leave --use_prompt_extend off.

Sources

Related posts

More in Models

All Models posts

Written by Sume