Document to video with Wan 3.0: what Sume accepts instead of a PDF
Alibaba's Wan 3.0 takes documents and webpages as references. Sume's catalog documents image, video and audio references, so convert pages to images first.

Wan 3.0 as Alibaba describes it can take documents, webpages and text as reference assets. On Sume, wan-3.0 is in the video catalog, and the Sume docs describe image, video and audio references, not document files, so you convert the pages to images first.
That is a small step and it is also the honest one. Do not send a PDF to a Sume video endpoint and expect it to be read. Export the pages, send them as image references, and say in the prompt what each page is.
Vendor claim versus the Sume catalog
Per the Wan 3.0 repository (read 2026-10-05), the model generates up to 30 seconds natively, accepts up to 20 reference assets including documents, webpages, text and images, and does instruction- and reference-based editing and automatic scene splitting. TechNode's launch report (read 2026-10-05) dates the launch to about 2026-08-24 and headlines the document input. The repository describes a closed beta and API, not open weights.
The Sume docs list wan-3.0 with a duration range of 2 to 30 seconds, and say Wan 3.0 accepts audio and video references as well as images. The docs do not list document or webpage inputs for it, so treat those as vendor features Sume does not document.
| Input | Alibaba Wan 3.0 repository | Sume docs for wan-3.0 |
|---|---|---|
| Duration | Native 30 s | 2 to 30 s |
| Reference assets | Up to 20: documents, webpages, text, images | Image, video and audio references; per-model limits in GET /v1/videos/models |
| Documents and webpages | Listed as reference types | Not listed; export pages as images |
| Text in video | Rendering in 12 languages | Prompt text only; check output before use |
| Editing | Instruction and reference based | Use the Video Router catalog capabilities for each model |
A practical conversion
Take a one-page product sheet or a short deck. The steps are yours, on your side, before you call Sume.
- Export each page as a PNG or JPEG. Public HTTPS URLs only: Sume rejects localhost, private-network and non-HTTPS URLs.
- Pick the three to six pages that carry the message. A reference per page is only useful if the page has one idea.
- Number them in the prompt in list order and say what each one shows: the price table, the product photo, the warranty line.
- Ask for a clip of 10 to 30 seconds that follows that order, and keep text in the prompt short and literal.
What a 30-second explainer costs
The prices come from the provider list times Sume's 1.25. At 720p Wan 3.0 is $0.125 per second, so a 15-second explainer is $1.875 and a 30-second one is $3.75. At 1080p it is $0.25 per second, so 30 seconds is $7.50. At 480p it is $0.0625 per second, so a 30-second draft is $1.875.
The limits of each model differ, so read supported_input_references and the capability fields from GET /v1/videos/models before you send twenty images and get a 400. Sume rejects parameters a model does not list rather than dropping them, which is easier to debug than a silent change.
Where it breaks
If the document has dense text, a generated video is the wrong format for the text. Render the key numbers and offer terms as captions or an end card, where you control the exact characters, and let the model animate the visuals. A model can paraphrase a price. A caption cannot.
A checklist before the first render
Here is the same advice as a checklist you can hand to a teammate before the first render. Confirm that the pages are public HTTPS URLs and that each one is a clean export. Confirm the model limits with the catalog call. Draft at 480p, where a 30-second clip is $1.875, and read the result before you pay for 1080p.
Then decide what the video is for. A product sheet becomes a 15-second teaser. A pricing page becomes a short explainer with the numbers in a caption. A slide deck for a webinar is usually better as a narrated slideshow than as generated motion, since the slides are already the visuals and the voice is the missing part. Sume's Timeline and text-to-speech endpoints cover that route, and the posts linked below price it.
Wan 3.0's document input is a vendor feature that Sume does not document, so build your plan on the image route. If Sume documents document inputs later, the plan gets simpler, and nothing you built will break.
Prepare the source first
Before you queue anything, turn the source into the form the API takes. For a slide deck, export each key slide as an image and keep the speaker notes as text. For a webpage, capture a clean screenshot of the section you want and copy the headline and offer into the prompt. For a document, write a short brief of three or four lines and attach any figures as images.
Then render one 5-second draft at the lowest resolution and check that the numbers and names on screen match your source. Generated video can alter small text, so anything that must be exact, such as a price, belongs in the caption or assembly step rather than in the generated frames.
Sources
Related posts
More in Use cases
- Does YouTube's AI label hurt reach or monetization? What it says
YouTube states a disclosure label alone does not change recommendations or monetization eligibility. What it covers, what it does not, and a workflow.
- Do resizing or color correction need Meta's AI info label?
Meta says minor edits such as resizing or color correction do not need an AI label, and not every ad shows one yet. What that means for a generated ad clip.
- Does YouTube's Shorts update demonetize AI videos?
The Oct 2026 change touches recommendations, not the Partner Program rules. How reach and monetization differ for AI Shorts, with YouTube's own wording.
- Draft at 480p, finish at 1080p: a Wan 3.0 holiday loop for $22.10
Thirty 5-second drafts at 480p ($9.60) plus five 10-second finals at 1080p ($12.50) cost $22.10 on Wan 3.0, against $37.50 for 30 clips straight at 1080p.
Written by Sume