HunyuanVideo 1.5 SSTA and FP8: or just call an API

HunyuanVideo-1.5 has 8.3B parameters, SSTA for a 1.87x speedup at 720p, FP8 GEMM and a 14GB VRAM floor. Decide whether to run it or call a hosted video API.

3 min readSume
All posts

HunyuanVideo-1.5 is an 8.3 billion parameter video model whose Selective Sliding Tile Attention (SSTA) gives a 1.87x speedup on 720p, and the project lists 14GB of VRAM as a minimum with model offloading enabled. If you do not want to operate GPUs, call a hosted model id instead and read its limits from the catalog.

What the repository states

The project page lists these facts.

HunyuanVideo-1.5 facts (read 2026-10-03)
ItemValue
Parameters8.3B
SSTA speedup at 720p1.87x
Clip length10-second 720p clips
Super-resolutionBuilt in, up to 1080p
PrecisionFP8 GEMM
Minimum VRAM14GB, with model offloading enabled

What running it costs you

Open weights save the per-second price but move work to you: driver and CUDA versions, FP8 kernel support on your GPU, queueing, storage for outputs, retries and monitoring. A 1.87x speedup is relative to the baseline attention, not to a hosted service, so it does not tell you your cost per clip.

A fair comparison needs your own numbers: GPU hours per clip at your target resolution, divided by how many clips a day you make, plus the time you spend keeping the stack alive.

The hosted alternative on Sume

Sume exposes a Video Router catalog at GET /v1/video-router/models and the OpenRouter-compatible GET /v1/videos/models. Each model reports supported resolutions, aspect ratios, durations and pricing_skus, and limits differ per model: for instance wan-3.0 accepts 2 to 30 seconds and most other catalog models are capped at 15 seconds.

I did not find a HunyuanVideo id in the Sume docs I read, so check the live catalog if you need that specific model. Billing is list price times 1.25 on every Video Router model, and jobs are asynchronous.

A quick decision rule

  • Run it yourself if you already own idle GPUs with at least 14GB (plus offloading) and need the weights.
  • Call a hosted id if clips are occasional or you need more than one model family.
  • Prototype on hosted first; self-host only if the volume justifies it.

Sources

Related posts

More in Models

All Models posts

Written by Sume