Open source AI video generator with audio: LTX-2 and MiniMax H3
Two open-weights video models make sound with the picture: LTX-2 and MiniMax H3. What each page says, and how a hosted API reports audio.

Yes: two open-weights video models generate audio together with the picture. Lightricks describes LTX-2 as generating synchronized video and audio within a single model, and MiniMax says its open-sourced H3 generates video with native stereo audio.
Read 2026-09-29: the LTX-2 model card and MiniMax's H3 open-source announcement. Sume's side comes from Video generation.
What do the two pages say about sound?
| Model | What the page says about audio |
|---|---|
| LTX-2 (Lightricks) | A DiT-based audio-video foundation model that generates synchronized video and audio within a single model; audio without speech may be of lower quality |
| MiniMax H3 | Generates video with native stereo audio; output audio is 32 kHz stereo; stable dialogue support for 11 languages |
Can I get generated audio from a hosted API instead?
Sume's video catalog reports audio per model through a generate_audio field, and the generate request accepts a generate_audio boolean that defaults to the model's audio capability. Two listed models are described with audio in the docs: minimax-h3-max with native stereo audio, and gemini-omni-flash-1.1 with native synced audio.
Read the field from the catalog rather than assuming it, because models differ. The docs do not say that the hosted MiniMax models run the open-sourced checkpoints, so do not treat the hosted ids and the downloadable weights as the same thing.
curl "https://api.sume.com/v1/catalog" \
-H "Authorization: Bearer $SUME_API_KEY"What input modes does the open MiniMax H3 release have?
MiniMax's page lists two base variants. H3-Base-FL2VA supports zero, one, or two input images (text-to-video, first-frame or last-frame, or first-and-last-frame). H3-Base-Ref2VA is an omni-reference mode with up to 9 images, up to 3 video clips and up to 3 audio clips, and audio must be accompanied by image or video input rather than used alone.
What audio can a hosted model take as input?
Audio in is a different feature from audio out. Sume's docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max, while Gemini Omni Flash 1.1 accepts video references but not audio. MiniMax's open-source page lists H3 output at 4-15 seconds and 24 FPS, so check the hosted catalog for each hosted id's own limits.
How do I add sound to a model that has none?
Video models that return silent clips can be paired with separate sound. AI video with sound: generate audio covers the options on Sume, and MiniMax H3 open weights on Hugging Face covers what the H3 release contains.
What does this answer leave out?
- Only cards read on 2026-09-29 count: other open models may also generate audio.
- Audio quality is not compared here; neither page gives a like-for-like measure.
- Both licenses are their own agreements (
ltx-2-community-license-agreementand MiniMax's H3 community license); read them before commercial use.
Sources
Related posts
More in Models
- Open-source video model vs API: which to use for Wan-class video
Weights you run yourself, or a hosted video API? What the choice changes for cost, setup and limits, with Wan 3.0's per-second API price as the worked example.
- Qwen Image 2.1 API: open weights, and what Sume lists
Qwen-Image-2.1 is an open-source 7B image model on Hugging Face with no inference provider. Sume lists qwen-image and qwen-image-max instead.
- Qwen Image license: is Qwen-Image open source under Apache 2.0?
The Qwen-Image model card says the weights are Apache 2.0 and were released on 2025-08-04. What that covers, and how the hosted qwen ids on Sume differ.
- Recraft V4 SVG: can you get vector output through Sume?
Recraft says V4 can make editable SVG. Sume's recraft/recraft-v4 advertises webp output only. What that means, plus a working call.
Written by Sume