Wan-Animate-2 Base vs Distillation weights: which to download
Wan-Animate-2 ships Base and Distillation checkpoints, each with its own YAML config. What the card says, and the hosted limits Sume lists for comparison.

The release notes list two checkpoints from August 07, 2026: Wan-Animate-2 Base and Wan-Animate-2 Distillation. Both come from the same Wan-AI/Wan2.2-Animate-2-14B download, and the card gives a separate command and YAML config for each. It does not state a speed or quality difference between them, so measure on your own clips.
Facts are from the model card, read 2026-10-01. Sume's hosted limits are from Video generation, read 2026-10-01.
What differs between the two commands?
Both run wan_animate_2_demo.py with --prompt, --refer-img-file and --refer-video-file. The config differs: ./wan_animate_2.yaml for Base and ./wan_animate_2_distillation.yaml for Distillation. The card also says to caption the reference image with an LLM first, and to adjust the parallel configs in the YAML files if your hardware is not 8 A800 GPUs.
| Checkpoint | Config file | Release note date |
|---|---|---|
| Base | ./wan_animate_2.yaml | August 07, 2026 |
| Distillation | ./wan_animate_2_distillation.yaml | August 07, 2026 |
Is Distillation the faster one?
The card does not say. The name suggests a distilled variant, but the page gives no step counts or timings for it, so treat that as an assumption. The card does describe a separate Wan-Animate-2-Lite variant aimed at real-time streaming, which is a different thing from the Distillation weights.
Do I need to download both?
No. The huggingface-cli download Wan-AI/Wan2.2-Animate-2-14B --local-dir ./ckpts/ command pulls the repo, and you pick a config at run time. Start with Base as the reference, then try Distillation on the same inputs and compare.
What do hosted limits look like for comparison?
Sume's catalog lists wan-3.0 at 2-30 seconds and a Motion Transfer model at 4-30 seconds; every other catalog model tops out at 15 seconds. A local run has no such catalog, so plan clip length yourself. For hosted submits, send Idempotency-Key on create so a retry returns the original job; see idempotency keys for AI video APIs.
Sources
Related posts
More in Models
- Which Sume video models accept an audio reference?
Seven Sume video models take reference audio: the four Seedance 2.x ids, Wan 3.0 and the two MiniMax H3 models. Kling, Grok and Gemini Omni Flash do not.
- Which Sume video models accept a reference video?
Nine Sume video ids take a video input: the four Seedance 2.x ids, Wan 3.0, both MiniMax H3 models, Omni Flash and Genjutsu. Kling and Grok do not.
- WorldCrafter camera control: Base weights, adapter and LoRA
TencentARC WorldCrafter-Base ships transformer weights with a matching camera adapter and LoRA. On Sume, steer camera motion in the video prompt instead.
- YouTube fhd, qhd, uhd thumbnails in the API: a 3840 AI image
YouTube's API notes fhd, qhd and uhd thumbnail keys for some videos. Sume's ChatGPT Image 2.5 image_size accepts a 3840 edge, enough for a 4K image.
Written by Sume