Wan2.2 S2V-14B speech-to-video: open weights vs Sume audio options
Wan2.2-S2V-14B is an Apache 2.0 audio-driven model at 480P and 720P. On Sume: an audio reference on Wan 3.0, or a script-driven Avatar video.
Wan2.2-S2V-14B is an open-weights speech-to-video model: you supply audio and it generates video driven by that audio, at 480P or 720P, under the Apache 2.0 license listed in the Wan2.2 repository (read 2026-10-02). If you want to avoid running a 14B model yourself, Sume has two hosted routes: pass an audio reference to Wan 3.0, or render a script-driven talking Avatar video.
What the Wan2.2 README lists
The README's model table (read 2026-10-02) names five Wan2.2 models. S2V-14B is the audio-driven one; its release is dated Aug 26, 2025 in the news list, the same list that dates Animate-14B to Sep 19, 2025.
| Model | Resolution | What it does |
|---|---|---|
| T2V-A14B | 480P and 720P | Text to video |
| I2V-A14B | 480P and 720P | Image to video |
| TI2V-5B | 720P | Text and image to video |
| S2V-14B | 480P and 720P | Audio-driven generation |
| Animate-14B | Variable | Character animation and replacement |
What running it yourself involves
The README says the 14B models want 80 GB or more of VRAM on a single GPU, or a multi-GPU setup with FSDP, while the 5B model needs at least 24 GB. So S2V-14B is a data-center workload, not a laptop one. You also own the audio preprocessing, queueing, storage and monitoring.
Route one: an audio reference on Wan 3.0
Sume's video docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Wan 3.0 (wan-3.0) accepts 2 to 30 seconds per clip. Pass an audio_url input reference with a public HTTPS link and a prompt describing the scene.
This is a reference, not a strict speech-to-video contract: the model uses the audio as guidance. Wan 3.0 is a different, hosted model from Wan2.2-S2V, so results will not match.
Route two: a script-driven Avatar video
If the goal is a person talking, Sume's Avatar video route takes a ready avatar and either a script or multi-scene video_inputs, for planned durations of 4 to 60 seconds at 720p. It is driven by a script, not by an audio file you upload, which is the key difference from S2V.
Which to choose
Sume does not host Wan2.2-S2V weights.
- Need a specific voice recording to drive the motion: S2V-14B, or an audio reference on Wan 3.0.
- Need a presenter reading a script: Avatar video.
- Need to control the weights and data path: self-host S2V and budget for the GPUs.
Sources
Related posts
More in Models
- Wan 3.0 allows 20,000 prompt characters, MiniMax H3 7,000
Alibaba allows 20,000 characters in a Wan 3.0 prompt; MiniMax caps each H3 text item at 7,000. How to write one prompt that fits both on Sume.
- Wan 3.0 input video plus output capped at 30 seconds: Sume plan
Alibaba caps Wan 3.0 input video plus output at 30 seconds. What that means for reference_video_urls and duration when you call wan-3.0 on Sume.
- Which Flow model can extend a video? Veo 3.1 Lite today, Omni soon
Flow's model table lets only Veo 3.1 Lite extend a clip, with an 8-second limit; Omni lists Extend as coming soon. Here is what to do on Sume meanwhile.
- Zalando designer minimum 1800x2600: request 1808x2608 on Sume
Zalando designer brands need at least 1800x2600 px in 1:1.44 JPEG. GPT custom sizes need multiples of 16, so ask Sume for 1808x2608, which clears both edges.
Written by Sume