Word document to video AI: turn a .docx into a narrated video
To turn a Word document into a video, extract its text and images first. Sume's agent takes text in input and images as attachments; a .docx isn't one.

A Word document becomes a video in three steps: get the text and pictures out of the .docx, write a short brief for the video, and send both to a video agent. On Sume, a run takes the text in input (up to 2 MiB) and up to 30 images as attachments; a .docx is not an attachment type, so extract it first.
Limits are from the Format API and Create a run docs, read 2026-09-29. Alibaba's Wan 3.0 API lists docx among the files it can parse (API reference); Sume's does not pass a document to a model.
Where does each part of the document go?
| Part of the .docx | Where it goes | Limit |
|---|---|---|
| Body text and headings | input, as JSON | 2 MiB, at most 64 top-level keys |
| Pictures and charts | attachments, type input_image | 30 images; JPEG, PNG, WebP, GIF, AVIF; 30 MB each |
| What to make | instruction | 8,000 accepted; about 4,000 reach the prompt |
How do I extract the text?
Use any docx reader in your own code, or save the file as plain text. Keep headings, drop headers, footers and page numbers, and put the result in a key you choose, such as input.text. Sume publishes no fixed field list for input; the agent reads it as caller data.
How do I send it?
One Agent Completion with a required generation_spend_cap_usd, then poll the run. PDF to video AI has the request body; a Word file goes in the same way once it is text.
What can go wrong with a Word file?
- Tracked changes and comments can end up in the extracted text; accept or strip them first.
- Tables flatten into lines; restate the numbers that matter in your brief.
- Embedded images are not extracted for you; export the ones you want and host them at public HTTPS URLs.
Sources
Related posts
More in Use cases
- Safety training videos for employees, made with AI
Make safety training videos for employees with an AI presenter: one hazard per clip, your own site photos, captions, and versions in your crew's languages.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
Written by Sume