Wan 3.0 thinking mode: why document and web page input needs it
fal says Wan 3.0 needs thinking mode to read a document or web page; Alibaba says prompt_extend must be true. What each means, and why Sume's wan-3.0 skips it.

Wan 3.0's thinking mode is the setting that lets it read a document or a public web page as input. fal's Wan 3 page says thinking is required for those two inputs; Alibaba's API reference names the matching parameter prompt_extend, which must be true whenever a file or a link is provided. Sume's wan-3.0 does not expose either input, so there is no thinking setting to turn on.
Read 2026-09-29 on fal's Wan 3 page and Alibaba Cloud's API reference; the Sume side is from the model catalog and Video generation.
What does each vendor call it?
The two pages use different words for what may be the same switch, and neither page says they are identical, so treat them as separate settings.
| Vendor page | Name | What it says |
|---|---|---|
| fal | Thinking | Required for the document and web-page reference inputs |
| Alibaba Cloud | prompt_extend | A language model rewrites the prompt; required when a document or web page is given; on by default |
What does prompt extension do to my prompt?
Alibaba says a language model rewrites the input prompt, that this significantly improves quality for shorter prompts, and that it increases latency. If you need the words you wrote to stay exactly as written, that is a trade to know about on any route that turns it on.
Is there a thinking switch on Sume?
No. The catalog constraints for wan-3.0 say file_url, web_url and enable_thinking are not exposed in v1. The fields you can send are prompt, duration, resolution, aspect_ratio, generate_audio, frame_images and input_references.
How do I write a good prompt without it?
Say who or what is in frame, what moves, how the camera moves, and the light, in one clear paragraph. That gives the model the content a document parse would otherwise supply. For a long source, summarize it into scenes yourself first.
Sources
Related posts
More in Models
- Wan 3.0 vertical video: 9:16 for Reels, Shorts and TikTok
Wan 3.0 supports 9:16. Send aspect_ratio 9:16 to wan-3.0 for a vertical clip of 2 to 30 seconds at 480p, 720p or 1080p. Accepted ratios and cost.
- Wan 3.0 vs MiniMax H3: clip length, ratios and reference caps
Wan 3.0 runs 2 to 30 s at up to 1080p; MiniMax H3 runs 5 to 15 s at 480p or 768p. Where they differ on Sume: ratios, reference caps and audio inputs.
- Wan 3.0 vs Seedance 2.0: which model to call when
Wan 3.0 and Seedance 2.0 are both on Sume's video API. Wan runs 2 to 30 s and bills per second; Seedance 2.0 runs 4 to 15 s with a 21:9 frame. Pick by the job.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume