Wan 3.0 document to video: what the model reads and what Sume passes

Alibaba's Wan 3.0 can read an uploaded document to make a video. On Sume's video API today wan-3.0 takes prompts, frames and media references, not a file input.

4 min readSume
All posts

Wan 3.0 can take a document as input on the vendor's own API, but Sume's wan-3.0 does not expose that field today. Alibaba Cloud describes a document-to-video feature that parses a file and generates a video from it; Sume's catalog lists wan-3.0 for text, first and last frames, and image, video and audio references, and notes that file_url, web_url and enable_thinking are not exposed in v1.

The Alibaba details are from its Wan3.0 guide and API reference; fal's Wan 3 page says documents need thinking mode. The Sume side is from Video generation and the current video catalog. All read 2026-09-29.

What can Wan 3.0 do with a document?

Per Alibaba, you provide one document or one public link and the model parses the content to build the video. The limits it publishes:

From Alibaba Cloud's API reference, read 2026-09-29.
ItemAlibaba's stated limit
File formatsdocx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md
File sizeNo more than 100 MB
File lengthNo more than 50 pages
Per requestOne file or one link, not both

Can I send a document to wan-3.0 on Sume?

No. The /v1/videos request has prompt, frame_images and input_references; a reference is an image_url, video_url or audio_url. There is no field for a document, so a PDF or a slide deck has nowhere to go on this route.

How do I get a video from a document on Sume, then?

Read the document yourself and turn it into a prompt. Extract the text, pick the three or four scenes worth showing, and write one prompt per scene; export the charts or pages you want on screen as images and pass them as input_references or a first_frame. PDF to video AI shows the agent route, where the extracted text goes in input.

Which route when?

  • You want a clip per scene, with control of each prompt: wan-3.0 on POST /v1/videos, 2 to 30 seconds each.
  • You want one call from a brief and material to a finished video: a Format or an Agent Completion, with the text in input.
  • You need the vendor's own document parsing: that is Alibaba's API, which Sume does not front today.

Sources

Related posts

More in Models

All Models posts

Written by Sume