PDF to video AI: how to turn a document into a video

PDF to video AI turns a document's text and figures into a narrated video. On Sume today, you extract the text and export the figures as images first.

5 min readSume
All posts

PDF to video AI turns a document into a narrated video: it reads the PDF's text and figures, writes a script, voices it, and pairs each line with a visual. With a video agent API such as Sume's, you do the reading step yourself today: extract the text, export the charts or pages you want shown as images, and send both with a brief, because a PDF can't be attached.

The Sume facts below come from the Agent Completions, Format API, and Create a run docs and the API reference, read on 2026-09-28.

Can I upload a PDF to an AI video agent?

Not as a file on Sume today. The Agent Completions docs list non-image attachments as not available yet: "input_image is the only type today; PDFs and other files follow later." The Format docs say to send documents by URL in input instead. They do not say how an agent reads that file, so send the extracted text as well, and the agent never has to parse the PDF.

From Agent Completions, Format API, and Create a run, read 2026-09-28.
Part of the PDFWhere it goesLimit
Body text, headings, key numbersinput (caller data)2 MiB, carried whole on a Format run
Charts, diagrams, page rendersattachments as input_image30 images; 30 MB each, 500 MB per run
The PDF itselfA URL inside input, per the Format docsNot an attachment type; how the agent reads it is not documented
What to make (length, shape, tone)instruction~4,000 characters reach the prompt on a Format run

How do I convert a PDF to video using AI?

The same request works for a report, a lesson, or a product guide. For a feed of articles rather than one document, Blog to video AI automates the call on each new post.

  • Extract the text with any PDF tool. Drop running headers, footers, and page numbers so they aren't read aloud.
  • Export each chart or page you want on screen as JPEG, PNG, WebP, GIF, or AVIF, and host it at a public HTTPS URL. Sume fetches each one when the run is created.
  • Write a short brief: length, aspect ratio, audience, and which figures must appear. How to write a video brief has a template.
  • Send one Agent Completion with the text in input, the images in attachments, and a generation_spend_cap_usd, which is required and has no default.
  • Poll the run until it is terminal. A completed run lists generated media in output.videos at durable media.sume.com URLs.
curl -sS -X POST "https://api.sume.com/v1/agent/completions" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "instruction": "Make a 60-second 16:9 explainer from the document text in the input. Use the attached charts. Use only facts from the document.",
    "input": { "title": "Q3 field report", "text": "…extracted PDF text…" },
    "attachments": [
      { "type": "input_image", "image_url": "https://example.com/chart-1.png" }
    ],
    "generation_spend_cap_usd": 5
  }'

How long will the video from a PDF be?

As long as the narration. A long PDF read in full makes a long video, so decide the length in the brief and let the script summarize. Text to speech time calculator shows how to estimate speech time from a word count.

If you build the video yourself instead, voice the script with TTS 1.0 (POST /v1/tts-1.0/generate). Its transcript takes up to 20,000 characters per request, and synthesized audio longer than 1200 s fails with tts_duration_exceeded. Split a long document into parts below both limits.

What does PDF to video AI not do?

Four limits to plan around:

  • It doesn't take the PDF as an attachment on Sume today, and there is no public route for uploading a local file.
  • It doesn't promise to redraw a chart exactly. If a figure has to be exact, attach it as an image and say in the brief that it must be shown as is.
  • It doesn't check the document. The script can only be as accurate as the text you send, so review the draft before you publish it.
  • It doesn't clear rights. Use documents and figures you're allowed to republish.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume