Wan2.2 S2V-14B speech-to-video: open weights vs Sume audio options

Wan2.2-S2V-14B is an Apache 2.0 audio-driven model at 480P and 720P. On Sume: an audio reference on Wan 3.0, or a script-driven Avatar video.

5 min readSume
All posts

Wan2.2-S2V-14B is an open-weights speech-to-video model: you supply audio and it generates video driven by that audio, at 480P or 720P, under the Apache 2.0 license listed in the Wan2.2 repository (read 2026-10-02). If you want to avoid running a 14B model yourself, Sume has two hosted routes: pass an audio reference to Wan 3.0, or render a script-driven talking Avatar video.

What the Wan2.2 README lists

The README's model table (read 2026-10-02) names five Wan2.2 models. S2V-14B is the audio-driven one; its release is dated Aug 26, 2025 in the news list, the same list that dates Animate-14B to Sep 19, 2025.

Wan2.2 open-weights models, vendor facts (read 2026-10-02)
ModelResolutionWhat it does
T2V-A14B480P and 720PText to video
I2V-A14B480P and 720PImage to video
TI2V-5B720PText and image to video
S2V-14B480P and 720PAudio-driven generation
Animate-14BVariableCharacter animation and replacement

What running it yourself involves

The README says the 14B models want 80 GB or more of VRAM on a single GPU, or a multi-GPU setup with FSDP, while the 5B model needs at least 24 GB. So S2V-14B is a data-center workload, not a laptop one. You also own the audio preprocessing, queueing, storage and monitoring.

Route one: an audio reference on Wan 3.0

Sume's video docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Wan 3.0 (wan-3.0) accepts 2 to 30 seconds per clip. Pass an audio_url input reference with a public HTTPS link and a prompt describing the scene.

This is a reference, not a strict speech-to-video contract: the model uses the audio as guidance. Wan 3.0 is a different, hosted model from Wan2.2-S2V, so results will not match.

Route two: a script-driven Avatar video

If the goal is a person talking, Sume's Avatar video route takes a ready avatar and either a script or multi-scene video_inputs, for planned durations of 4 to 60 seconds at 720p. It is driven by a script, not by an audio file you upload, which is the key difference from S2V.

Which to choose

Sume does not host Wan2.2-S2V weights.

  • Need a specific voice recording to drive the motion: S2V-14B, or an audio reference on Wan 3.0.
  • Need a presenter reading a script: Avatar video.
  • Need to control the weights and data path: self-host S2V and budget for the GPUs.

Sources

Related posts

More in Models

All Models posts

Written by Sume