Mac or Windows team: which open video models need Linux and CUDA?
HunyuanVideo and HunyuanVideo-1.5 list Linux; LTX-2.5's fastest decoder is Linux plus CUDA; H3 has a Mac route. What each page says and a hosted option.

Two of the four open video models name Linux outright: the HunyuanVideo README lists Linux as the tested operating system, and the HunyuanVideo-1.5 model card lists Linux as the operating system with CUDA. LTX-2.5's fastest VAE backend is Linux plus CUDA only, with a slower fallback on Windows and macOS. MiniMax H3's page lists Apple Silicon routes. Wan 2.2's README gives no operating-system line in the sections I read. Vendor pages read 2026-10-08.
What each page says
Short table of the operating-system statements, quoted or paraphrased from the vendor pages.
| Model | Statement | Source page |
|---|---|---|
| HunyuanVideo | Tested operating system: Linux; NVIDIA GPU with CUDA required | README system requirements |
| HunyuanVideo-1.5 | Operating system: Linux; Python 3.10 or higher; minimum 14 GB GPU memory with offloading | Hugging Face model card |
| LTX-2.5 | natten extra is Linux + CUDA only; on Windows and macOS it is skipped and decoding falls back to Triton or eager | LTX-2 README quick start |
| MiniMax H3 | Apple Silicon works via MLX (128 GB unified memory recommended) or an NF4 route from 16 GB; Colab GPU runtimes work | H3 FAQ |
| Wan 2.2 | No operating-system line found in the README sections read; commands use CUDA tools such as torchrun | README |
What this means for a laptop team
Linux and CUDA are a reasonable requirement for a render server and a poor one for a designer's MacBook. The practical split is: develop prompts anywhere, run models on a Linux GPU box. For H3, the vendor page names an Apple Silicon path but with large memory, which is a workstation purchase.
None of these pages promise parity between the Linux path and the fallback. The LTX README says only that the fallback exists and is slower than the natten backend.
A hosted job runs from any OS
A hosted job needs only an HTTPS client. Sume's video route is POST https://api.sume.com/v1/videos with a bearer key; the response gives a job id and a polling URL, and GET /v1/videos/{jobId} reports pending, in_progress, completed, failed or cancelled. A Mac, a Windows laptop and a Linux CI runner behave the same.
The cost is per second: $0.375 for five seconds at 768p on minimax-h3, or $0.625 on wan-3.0 at 720p (list x 1.25). Sume does not list HunyuanVideo or LTX.
Decision
Pick by where the work runs.
- Own a Linux GPU box and want weights: Hunyuan, LTX or Wan 2.2 (mind each license).
- Mac only: use a hosted id, or H3 with a large unified-memory Mac.
- Windows laptop for design, Linux server for batches: hosted jobs keep the laptop out of the loop.
A note on the sources
These are statements on vendor pages as of 2026-10-08. A page that lists Linux as the tested system does not forbid other systems; it says what was tested. Community ports to Windows and macOS exist for several of these models, but this post cites only vendor pages, so it makes no claim about their quality.
Sources
- HunyuanVideo README, Tencent-Hunyuan on GitHub (read 2026-10-08)
- HunyuanVideo-1.5 model card, Hugging Face (read 2026-10-08)
- LTX-2 README, Lightricks on GitHub (read 2026-10-08)
- MiniMax H3 Open: FAQ and deployment, MiniMax (read 2026-10-08)
- Wan2.2 README, Wan-Video on GitHub (read 2026-10-08)
- Sume docs: Video generation
Related posts
More in Models
- MAI-Transcribe-2: 60 languages, 17 more than 1.5, and the Sume hint
Microsoft lists 60 languages for MAI-Transcribe-2, 43 for MAI-Transcribe-1.5, so 17 new ones. Sume STT takes a language_code hint or auto-detects.
- MAI-Transcribe-2-Streaming runs in 4 regions, MAI-Voice in 14
Microsoft lists four regions for MAI-Transcribe-2-Streaming and 14 for MAI-Voice-2.1; three overlap. Sume's STT and TTS requests have no region field.
- MAI-Voice-2.1 languages: 23 listed, no Japanese or Arabic voice
Microsoft's MAI-Voice-2.1 voice table covers 23 languages and 28 locale codes, with no Japanese or Arabic row. How Sume's TTS language field handles both.
- MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.
Written by Sume