HunyuanVideo 1.5 cache inference: DeepCache, TeaCache, TaylorCache

HunyuanVideo 1.5 added DeepCache on Nov 24 2025 and TeaCache plus TaylorCache on Nov 27, switched with --enable_cache and --cache_type. What the README claims.

4 min readSume
All posts

HunyuanVideo 1.5 supports three cache methods in its own code: DeepCache, added on Nov 24, 2025, and TeaCache and TaylorCache, added on Nov 27, 2025. You turn them on with --enable_cache and choose one with --cache_type. The README gives no number for the three built-in methods, only the comment that the cache "significantly speeds up inference". The 1.7x at 20 steps it mentions belongs to a separate ComfyUI-MagCache integration.

What does the README list, and when?

HunyuanVideo-1.5 speed-related changes (README news, read 2026-10-09)
DateChange
Nov 20, 2025initial inference code and weights
Nov 24, 2025DeepCache inference
Nov 27, 2025cache inference: deepcache, teacache, taylorcache
Dec 05, 2025480p I2V step-distilled model, 8 or 12 steps
Dec 23, 2025fp8 gemm inference

Can the speedups be stacked?

The README example command passes the flags together: --enable_step_distill, --sparse_attn, --use_sageattn, --enable_cache and --cache_type, --overlap_group_offloading. The argument table lists --enable_cache as default false and --cache_type as default deepcache, with --cache_start_step (11), --cache_end_step (45) and --cache_step_interval (4) to tune which steps are skipped. It does not say which combinations are supported, or what each costs in quality. Change one flag at a time, and keep each output with the flags that made it.

A cache trades some quality for speed in general; the README does not quantify the trade for HunyuanVideo, so measure it on your own prompts.

How do you test a cache setting?

Fix the prompt, the image and the seed. Run the model with caching off and note the time and the output. Then run each of the three cache types in turn with everything else unchanged. Compare the frames side by side at full size: caches tend to show up as softer detail or small motion artifacts. The README does not describe what to expect, so this comparison is the only evidence you will have.

Record the step count too. A cache that helps at 50 steps may give less at 8 to 12 steps with the step-distilled file, since there are fewer repeated steps to skip.

Why does this matter for a cost estimate?

A faster run lowers GPU-seconds per clip, which lowers the break-even against per-second pricing. A hosted comparison, if you need one, is per second of video: Sume lists models such as wan-3.0 (2 to 30 seconds) and minimax-h3 (5 to 15 seconds); the video docs explain how to read the price from the catalog. Sume does not list HunyuanVideo (catalog code, read 2026-10-09).

Sources

Related posts

More in Developers

All Developers posts

Written by Sume