HunyuanVideo 1.5 cache inference: DeepCache, TeaCache, TaylorCache
HunyuanVideo 1.5 added DeepCache on Nov 24 2025 and TeaCache plus TaylorCache on Nov 27, switched with --enable_cache and --cache_type. What the README claims.

HunyuanVideo 1.5 supports three cache methods in its own code: DeepCache, added on Nov 24, 2025, and TeaCache and TaylorCache, added on Nov 27, 2025. You turn them on with --enable_cache and choose one with --cache_type. The README gives no number for the three built-in methods, only the comment that the cache "significantly speeds up inference". The 1.7x at 20 steps it mentions belongs to a separate ComfyUI-MagCache integration.
What does the README list, and when?
| Date | Change |
|---|---|
| Nov 20, 2025 | initial inference code and weights |
| Nov 24, 2025 | DeepCache inference |
| Nov 27, 2025 | cache inference: deepcache, teacache, taylorcache |
| Dec 05, 2025 | 480p I2V step-distilled model, 8 or 12 steps |
| Dec 23, 2025 | fp8 gemm inference |
Can the speedups be stacked?
The README example command passes the flags together: --enable_step_distill, --sparse_attn, --use_sageattn, --enable_cache and --cache_type, --overlap_group_offloading. The argument table lists --enable_cache as default false and --cache_type as default deepcache, with --cache_start_step (11), --cache_end_step (45) and --cache_step_interval (4) to tune which steps are skipped. It does not say which combinations are supported, or what each costs in quality. Change one flag at a time, and keep each output with the flags that made it.
A cache trades some quality for speed in general; the README does not quantify the trade for HunyuanVideo, so measure it on your own prompts.
How do you test a cache setting?
Fix the prompt, the image and the seed. Run the model with caching off and note the time and the output. Then run each of the three cache types in turn with everything else unchanged. Compare the frames side by side at full size: caches tend to show up as softer detail or small motion artifacts. The README does not describe what to expect, so this comparison is the only evidence you will have.
Record the step count too. A cache that helps at 50 steps may give less at 8 to 12 steps with the step-distilled file, since there are fewer repeated steps to skip.
Why does this matter for a cost estimate?
A faster run lowers GPU-seconds per clip, which lowers the break-even against per-second pricing. A hosted comparison, if you need one, is per second of video: Sume lists models such as wan-3.0 (2 to 30 seconds) and minimax-h3 (5 to 15 seconds); the video docs explain how to read the price from the catalog. Sume does not list HunyuanVideo (catalog code, read 2026-10-09).
Sources
Related posts
More in Developers
- Hy Image 3.5 multi-turn editing with assembled_history vs Sume edits
How Tencent's Hy Image 3.5 Preview chains edits with assembled_history, and what the same loop looks like on Sume's stateless images route.
- Idempotency key from the order id, not a fresh UUID per attempt
A random UUID generated inside the retry loop gives every attempt a new key and a new paid job. Derive the Sume Idempotency-Key from the order instead.
- Inspect transcript words without start or end: skip them in cuts
Inspect transcript words can omit start or end. Treat an untimed word as unknown, and never cut across a gap that contains one: 0 of its time is safe to remove.
- IVR phone menu prompts with TTS: 8 kHz mu-law WAV and cost per prompt
Make phone-menu prompts as 8 kHz mu-law or A-law WAV with Sume TTS. A 20-prompt menu bills 20 cents because each short job hits the 1-cent minimum.
Written by Sume