PDF to video AI: how to turn a document into a video
PDF to video AI turns a document's text and figures into a narrated video. On Sume today, you extract the text and export the figures as images first.

PDF to video AI turns a document into a narrated video: it reads the PDF's text and figures, writes a script, voices it, and pairs each line with a visual. With a video agent API such as Sume's, you do the reading step yourself today: extract the text, export the charts or pages you want shown as images, and send both with a brief, because a PDF can't be attached.
The Sume facts below come from the Agent Completions, Format API, and Create a run docs and the API reference, read on 2026-09-28.
Can I upload a PDF to an AI video agent?
Not as a file on Sume today. The Agent Completions docs list non-image attachments as not available yet: "input_image is the only type today; PDFs and other files follow later." The Format docs say to send documents by URL in input instead. They do not say how an agent reads that file, so send the extracted text as well, and the agent never has to parse the PDF.
| Part of the PDF | Where it goes | Limit |
|---|---|---|
| Body text, headings, key numbers | input (caller data) | 2 MiB, carried whole on a Format run |
| Charts, diagrams, page renders | attachments as input_image | 30 images; 30 MB each, 500 MB per run |
| The PDF itself | A URL inside input, per the Format docs | Not an attachment type; how the agent reads it is not documented |
| What to make (length, shape, tone) | instruction | ~4,000 characters reach the prompt on a Format run |
How do I convert a PDF to video using AI?
The same request works for a report, a lesson, or a product guide. For a feed of articles rather than one document, Blog to video AI automates the call on each new post.
- Extract the text with any PDF tool. Drop running headers, footers, and page numbers so they aren't read aloud.
- Export each chart or page you want on screen as JPEG, PNG, WebP, GIF, or AVIF, and host it at a public HTTPS URL. Sume fetches each one when the run is created.
- Write a short brief: length, aspect ratio, audience, and which figures must appear. How to write a video brief has a template.
- Send one Agent Completion with the text in
input, the images inattachments, and ageneration_spend_cap_usd, which is required and has no default. - Poll the run until it is terminal. A completed run lists generated media in
output.videosat durablemedia.sume.comURLs.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"instruction": "Make a 60-second 16:9 explainer from the document text in the input. Use the attached charts. Use only facts from the document.",
"input": { "title": "Q3 field report", "text": "…extracted PDF text…" },
"attachments": [
{ "type": "input_image", "image_url": "https://example.com/chart-1.png" }
],
"generation_spend_cap_usd": 5
}'How long will the video from a PDF be?
As long as the narration. A long PDF read in full makes a long video, so decide the length in the brief and let the script summarize. Text to speech time calculator shows how to estimate speech time from a word count.
If you build the video yourself instead, voice the script with TTS 1.0 (POST /v1/tts-1.0/generate). Its transcript takes up to 20,000 characters per request, and synthesized audio longer than 1200 s fails with tts_duration_exceeded. Split a long document into parts below both limits.
What does PDF to video AI not do?
Four limits to plan around:
- It doesn't take the PDF as an attachment on Sume today, and there is no public route for uploading a local file.
- It doesn't promise to redraw a chart exactly. If a figure has to be exact, attach it as an image and say in the brief that it must be shown as is.
- It doesn't check the document. The script can only be as accurate as the text you send, so review the draft before you publish it.
- It doesn't clear rights. Use documents and figures you're allowed to republish.
Sources
Related posts
More in Agents
- Remote MCP server URL: what it is and where to find it
A remote MCP server URL is the HTTPS address of a server's MCP endpoint. Where to get one, where to paste it, and why it isn't a page to open.
- Video shot list template: columns, example, AI version
A video shot list is one row per shot: number, what we see, framing, move, length, audio, source. A template to copy, plus the columns AI video needs.
- What is an MCP server? A plain definition with examples
An MCP server is a program that gives AI apps tools, data, and prompt templates through the Model Context Protocol. How it works, with examples.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
Written by Sume