Wan 3.0 takes documents as references. What Sume's route accepts

Alibaba lists documents, webpages and text among Wan 3.0's 20 reference assets. Sume's video route takes image, video and audio references. How to bridge it.

5 min readSume
All posts

Alibaba's Wan 3.0 README (read 2026-10-05) lists up to 20 reference assets and names documents, webpages, text and images among them. On Sume, the video request documents three reference types: image_url, video_url and audio_url in supported_input_references. So a PDF or a URL to a webpage is not a documented input on Sume's route. The practical bridge is to turn the document into the things the model can use: a written prompt and a few images.

What each side documents

The Sume claim comes from one sentence in the video docs: the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max accept audio and video references, and a model accepts a reference type only if its supported_input_references lists it.

Reference inputs for Wan 3.0, vendor versus Sume docs, read 2026-10-05
InputWan 3.0 READMESume `wan-3.0` route
ImagesYes, within 20 assetsImage references via input_references
DocumentsYesNot listed
WebpagesYesNot listed
TextYesThe prompt field
Video / audioNot the focus of the README's reference listWan 3.0 accepts audio and video references per the Sume video docs

Turning a brief into inputs Sume takes

  • Pull the facts that matter from the document into the prompt, in the order the shots should appear.
  • Export the two or three pages or product pictures that carry the look as PNGs and pass them as image references.
  • Keep the prompt concrete: subject, action, camera, setting, one line of sound. Long pasted text is not the same as a brief.
  • Check the model's catalog entry for its real maximum before you send a large reference list.

Request shape

Image references go in an input_references array of image_url objects, as in the documented example. Wan 3.0 takes 2 to 30 seconds on Sume.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wan-doc-brief-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Product explainer: pan across the device, then a close-up of the dial. Calm, clean studio light.",
    "input_references": [
      { "type": "image_url", "image_url": { "url": "https://example.com/page-3.png" } }
    ],
    "duration": 20,
    "resolution": "720p"
  }'

What gets lost in the translation

Summarizing a document into a prompt discards its layout and ordering cues. Two habits recover most of that. First, write the prompt as a shot list, with one sentence per shot, in the order of the document's sections. Second, export one image per section that represents it, and name the files in order so that the reference list follows the document. You lose the model reading the original page, but you keep the structure.

If the original document includes charts or UI screens, treat those as images rather than text: a clean screenshot is a better reference than a paragraph describing it. Remove anything you would not want echoed in the video, such as confidential figures or internal names.

  • One shot per document section keeps the video readable.
  • Screenshots with clear subjects work better than dense slides.
  • Check the result against the document, line by line, before delivery.

Where to put the effort

Spend it on the prompt and on image selection. Those are the two inputs the hosted route does take, and they decide most of the result. A tight shot list of five lines with three clean images usually beats a pasted page of text.

Check the Wan 3.0 README for the vendor's own wording, and Sume's video docs for the fields your request must use. If a document-as-reference workflow is central to your job, say so in your evaluation notes: it is a gap between the vendor surface and the hosted route, not something to paper over.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume