Animate an infographic with Wan 3.0: first frame or reference image?
Use the infographic as a first frame to keep the layout, or as a reference to get new motion. How Sume decides, and which to pick for charts, icons and text.

To animate an infographic with Wan 3.0 and keep its layout, send it as a first frame (frame_images with frame_type: "first_frame" on /v1/videos, or image_url on the Video Router). To get new motion that only borrows its look, send it as a reference (input_references or reference_image_urls). Sume's Video generation docs define the two modes: frame images start an image-to-video job, references give the model visual guidance but not exact frames, and if you send both fields, frame_images controls the mode. Alibaba's Wan 3.0 README (read 2026-10-05) lists up to 20 reference assets and text rendering in 12 languages; read both as upper bounds, not promises for one chart.
What each mode does to an infographic
As a first frame, the infographic is frame zero: the headline, icons and numbers are exactly what you drew. The model then animates forward, and over a few seconds the layout can melt, because nothing holds the text in place after the first frame. As a reference, the infographic is only guidance, so the clip will not contain your layout at all, only a similar palette and style.
| Goal | Use | Field | Expect |
|---|---|---|---|
| Layout visible at start, then motion | First frame | frame_images or image_url | Exact frame 0; later frames may drift |
| Land on a given final layout | First and last frame | frame_images with first_frame and last_frame | A morph between two designed states |
| New scene in the same style | Reference | input_references or reference_image_urls | Style carried, layout not |
| Many icons as ingredients | References, up to 10 | reference_image_urls | Icons reused as objects in a scene |
Keep the clip short
The longer the clip, the more the text can drift. For a layout-faithful animation, ask for a short duration, such as 4 to 6 seconds, and use a first frame plus a last frame designed in your own tool, so the model only has to move between two states you made. The docs list end_frame support for Wan 3.0 in the catalog, and the Video Router's image-to-video fields are image_url and end_image_url.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: infographic-001" \
-d '{
"model": "wan-3.0",
"prompt": "Bars grow upward one by one, icons pulse once, camera static.",
"image_url": "https://example.com/infographic-start.png",
"end_image_url": "https://example.com/infographic-end.png",
"resolution": "720p",
"duration": 5,
"mode": "async"
}'When the text must be exact
Do not trust motion with small type. Animate the bars, icons and background, then lay the real labels over the result. Timeline compose does one still plus one video, and video captions burns authored cues. Put big numbers in captions, keep legends in the still, and treat anything the model draws as decoration.
Keep the prompt short and about motion only, since the picture already says what the content is. A practical order of work: design the start and end frames as ordinary images in your design tool, export them at the same size and the same aspect ratio as the output you will request, host them at stable https URLs, and generate a 480p draft first. The draft shows you whether the in-between motion is acceptable before you pay for the larger size. If the motion keeps breaking the layout, split the animation into two five-second clips, each between its own pair of frames, and join them.
Check the last frame with a still before you ship. If the last frame no longer matches the data, shorten the clip or replace the model's frames with your own final card using Timeline 1.0.
Sources
Related posts
More in Media tools
- Assert an edited image has the source's pixel size: Python check
A short Python check that downloads the source and the edited file and fails if the pixel size differs. Needed on Sume whenever you send an aspect_ratio.
- Audio detach range without end: take the rest of a track
In Sume audio detach, range.end is optional: a start-only range runs to the end of the track. How it meets the 900 s output cap and 1800 s source cap.
- audio_source_in_requires_single_spine: source_in needs one voice file
audio.source_in only works with one audio.url. With audio.parts, set source_in on each part instead; the Sume hint is use_part_source_in.
- Avatar clip frame rate: Griffin 25 fps vs Fabric vs H3 Max
Tavus says Griffin streams 8-frame latents at 25 fps. Sume's docs say Fabric is 25 fps and H3 Max must be measured. Probe a clip with video-inspect.
Written by Sume