Paper abstract to a 30 s explainer: Wan 3.0 scenes, claims as captions

Turn a research abstract into a 30 s Wan 3.0 explainer on Sume without invented results: the model draws the scene, you type every claim as a caption.

5 min readSume
All posts

A research abstract can become a 30 second explainer if the video shows the setting and your own captions state the findings. On Sume, ask wan-3.0 for a clip of up to 30 seconds from a prompt and up to 10 reference images, then burn the claims, numbers and citation with video captions cues. Alibaba's Wan 3.0 README (read 2026-10-05) lists document input and native 30 second generation; Sume does not expose a document field for wan-3.0, so you read the abstract yourself and decide what the pictures should be. That split is a good one: the model is bad at facts, and an abstract is nothing but facts.

Split the abstract three ways

Read the abstract and mark every sentence as setting, method or result. Only the first two can be pictured. Results are numbers and relationships, and they go into captions.

  • Setting: what is studied, as a scene (a coral reef, a data center, a classroom).
  • Method: one visible action (sampling, measuring, training, comparing).
  • Result: not pictured. Typed as a caption, with the exact figure and unit from the paper.
  • Limits and caveats: also typed, in one short line, so the video does not overstate.
  • Citation: authors, venue and year as the closing caption.

Why figures are a risky reference

It is tempting to send the paper's figures as reference images. Charts carry small axis labels and legends, and a video model will treat them as texture and may redraw them with different numbers. If a chart matters, show the original image untouched with Timeline compose, which puts one still and one video on screen together, and keep the generated footage for the scene around it. Use figures as reference_image_urls only for style, such as a color palette.

Which part of the explainer comes from where (read 2026-10-05)
PartSourceTool
Scene and moodYour prompt, 1 to 3 style referenceswan-3.0, up to 30 s, up to 10 images
Findings and numbersTyped from the paperVideo captions cues with text, start, end
The paper's own figureThe original imageTimeline compose, still plus video
Citation lineTyped from the paperVideo captions cue at the end

The generation request

The request is a plain Video Router job. Keep the prompt free of numbers and claims.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: paper-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A calm explainer scene: researchers at a lab bench sample water from small tanks, then a screen shows slow abstract graphs rising. Soft daylight, no readable text, no logos.",
    "reference_image_urls": ["https://example.com/palette-ref.png"],
    "resolution": "720p",
    "duration": 24,
    "aspect_ratio": "16:9",
    "mode": "async"
  }'

Captions, and a review step

Send the finished clip to video captions with a cues list. Each cue is text, start and end in seconds. Put the result at about four seconds, the caveat right after it, and the citation last.

Before you publish, ask the paper's author or a colleague who knows the field to read the captions against the paper. The video should never state more than the abstract does, and you should not name a person or institution in the footage unless you have their agreement.

Plan the timing from the captions, not from the clip. Thirty seconds holds about four short caption cards if each stays up for five seconds and the first one waits until the scene has settled. Write the cards first, then describe a scene whose beats match them. If you are summarizing several papers, make one clip per paper and join them with Timeline 1.0 rather than packing two abstracts into one prompt, because the model has no way to keep two sets of findings apart.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume