Wan 3.0 reads PDFs: what Sume's video request accepts
Wan 3.0 adds document and webpage inputs. Sume's video request takes prompt, image, video and audio fields only, so turn a PDF into a prompt first.

Wan 3.0 adds document and webpage inputs, covering PDFs, PPTs, Word documents and Excel sheets, according to its launch coverage. Sume's video request does not accept documents: it takes prompt, frame_images, input_references and related media fields. Summarize the document into a prompt yourself, then send it to wan-3.0.
What the launch post describes
The launch post also says Wan 3.0 recommends a video length from the prompt and includes an extension tool to lengthen videos. Those are model-side features as reported there. On Sume, you choose duration yourself, and wan-3.0 accepts 2 to 30 seconds.
What Sume's video request takes
The documented request fields are narrow, and this table lists them.
| Field | Type | Use |
|---|---|---|
| prompt | string | Text description, required |
| duration | integer | Length in seconds |
| frame_images | array | First and last frame images |
| input_references | array | Reference images, plus video and audio on models that accept them |
| resolution, aspect_ratio | string | Output shape |
From a PDF to a clip
Wan 3.0 honors audio and video references on Sume, but none of these fields carries a PDF. Turn the document into text first. A short pipeline works well:
- Extract the text from the PDF with your own tool
- Summarize it into a scene description under a few sentences
- Pull one or two key images from the document and host them at public HTTPS URLs
- Send the summary as
promptand the images asinput_references
Keep a fact check
Keep the document's facts in your own check step. A prompt that paraphrases a spec sheet can drift from the source, so review the output against the document before it goes out.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: doc-to-video-001" \
-d '{"model":"wan-3.0","prompt":"Animated explainer of a three-step product setup, clean flat style","duration":12}'Sources
Related posts
More in Developers
- Fetch tool to media input: which URLs Sume accepts
A model can fetch a page, but Sume media inputs must be public HTTPS image or video URLs, not web pages. Here is what each Sume endpoint accepts.
- What a media MCP server should declare at server/discover
MCP 2026-07-28 adds a required server/discover call. A media server has more to say than versions: async jobs, wait limits, scopes. Where Sume documents each.
- Which Lyria ran? job.model vs job.request.routed_model on Sume
On Sume's Music Router, job.model echoes what you sent and job.request.routed_model names the engine that ran. How to read both and pin a Lyria id.
- whisper-1 or gpt-transcribe for subtitles: what OpenAI assigns to each
OpenAI recommends gpt-transcribe, gpt-4o-transcribe-diarize for speakers, whisper-1 for translation and subtitles. Plus the 25 MB limit and a chunking script.
Written by Sume