Paper abstract to a 30 s explainer: Wan 3.0 scenes, claims as captions
Turn a research abstract into a 30 s Wan 3.0 explainer on Sume without invented results: the model draws the scene, you type every claim as a caption.

A research abstract can become a 30 second explainer if the video shows the setting and your own captions state the findings. On Sume, ask wan-3.0 for a clip of up to 30 seconds from a prompt and up to 10 reference images, then burn the claims, numbers and citation with video captions cues. Alibaba's Wan 3.0 README (read 2026-10-05) lists document input and native 30 second generation; Sume does not expose a document field for wan-3.0, so you read the abstract yourself and decide what the pictures should be. That split is a good one: the model is bad at facts, and an abstract is nothing but facts.
Split the abstract three ways
Read the abstract and mark every sentence as setting, method or result. Only the first two can be pictured. Results are numbers and relationships, and they go into captions.
- Setting: what is studied, as a scene (a coral reef, a data center, a classroom).
- Method: one visible action (sampling, measuring, training, comparing).
- Result: not pictured. Typed as a caption, with the exact figure and unit from the paper.
- Limits and caveats: also typed, in one short line, so the video does not overstate.
- Citation: authors, venue and year as the closing caption.
Why figures are a risky reference
It is tempting to send the paper's figures as reference images. Charts carry small axis labels and legends, and a video model will treat them as texture and may redraw them with different numbers. If a chart matters, show the original image untouched with Timeline compose, which puts one still and one video on screen together, and keep the generated footage for the scene around it. Use figures as reference_image_urls only for style, such as a color palette.
| Part | Source | Tool |
|---|---|---|
| Scene and mood | Your prompt, 1 to 3 style references | wan-3.0, up to 30 s, up to 10 images |
| Findings and numbers | Typed from the paper | Video captions cues with text, start, end |
| The paper's own figure | The original image | Timeline compose, still plus video |
| Citation line | Typed from the paper | Video captions cue at the end |
The generation request
The request is a plain Video Router job. Keep the prompt free of numbers and claims.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: paper-001" \
-d '{
"model": "wan-3.0",
"prompt": "A calm explainer scene: researchers at a lab bench sample water from small tanks, then a screen shows slow abstract graphs rising. Soft daylight, no readable text, no logos.",
"reference_image_urls": ["https://example.com/palette-ref.png"],
"resolution": "720p",
"duration": 24,
"aspect_ratio": "16:9",
"mode": "async"
}'Captions, and a review step
Send the finished clip to video captions with a cues list. Each cue is text, start and end in seconds. Put the result at about four seconds, the caveat right after it, and the citation last.
Before you publish, ask the paper's author or a colleague who knows the field to read the captions against the paper. The video should never state more than the abstract does, and you should not name a person or institution in the footage unless you have their agreement.
Plan the timing from the captions, not from the clip. Thirty seconds holds about four short caption cards if each stays up for five seconds and the first one waits until the scene has settled. Write the cards first, then describe a scene whose beats match them. If you are summarizing several papers, make one clip per paper and join them with Timeline 1.0 rather than packing two abstracts into one prompt, because the model has no way to keep two sets of findings apart.
Sources
Related posts
More in Use cases
- Failed payment reminder video: 20-second AI avatar script and cost
A friendly failed-payment clip from a Sume avatar: script rules, no scare wording, the API request and the cost per account on standard, plus and max.
- Same-second thumbnail frames for every episode of a Shorts series
YouTube lets each Short in a series have a custom thumbnail. Pull one frame per episode at the same second with Sume video-frames so the set looks consistent.
- Performance Max: 15 vertical + 15 horizontal Omni clips, $56.40
Performance Max takes 15 videos per orientation and one vertical 10-60 s clip for Shorts. Thirty 10 s Omni 1080p clips cost $56.40 on Sume.
- Performance Max builds a video if you upload none: $1.88 swap
Google makes Performance Max videos from images and text only if you upload none. A 10-second vertical Omni 1080p clip is $1.88 on Sume; you set the AI label.
Written by Sume