HunyuanVideo 720p: 1,904 s on one GPU, 338 s on eight

README latency for a 1280x720, 129-frame, 50-step clip: 1,904 s on 1 GPU and 338 s on 8. The GPU-seconds each setup burns, and what it means for hosted jobs.

5 min readSume
All posts

Tencent's HunyuanVideo README reports latency for a 1280x720 clip of 129 frames at 50 steps: 1,904.08 seconds on one GPU, 934.09 on two, 514.08 on four and 337.58 on eight. That is a 5.64x speedup on eight GPUs, but the eight-GPU run burns 2,700.64 GPU-seconds against 1,904.08 for one. Faster wall-clock time costs more total GPU time. Numbers from the README, read 2026-10-08; the GPU-seconds column is multiplication.

The table

The README gives the first three data columns (GPUs, latency, speedup). The last column is latency times the GPU count, which is what you pay for if you rent by the GPU-second.

HunyuanVideo 1280x720, 129 frames, 50 steps, README read 2026-10-08
GPUsLatencySpeedup vs 1 GPUGPU-seconds (latency x GPUs)
11904.08 s1.00x1,904.08
2934.09 s2.04x1,868.18
4514.08 s3.70x2,056.32
8337.58 s5.64x2,700.64

How to read it for a render queue

If you need one clip as fast as possible, use eight GPUs: about 5.6 minutes. If you need many clips and throughput matters more than latency, one GPU per clip is the cheapest per clip, because parallel efficiency falls as GPUs are added. By my arithmetic from the README numbers, efficiency (1,904.08 divided by GPU-seconds) is about 102 percent at two GPUs, 93 percent at four and 71 percent at eight.

Memory still applies: the same README says 60 GB is the minimum for 720x1280 at 129 frames, so 'one GPU' means a large-memory GPU.

What a hosted job changes

On a hosted API the batching, GPU count and queue are the provider's problem. Sume's video flow returns a job id at once and you poll GET /v1/videos/{jobId} or use a webhook. Sume's best-practices note says generation usually takes from 30 seconds to several minutes depending on model and parameters, and Sume's admission docs describe queued states when capacity or balance is short.

Sume does not list HunyuanVideo, so this is not a like-for-like swap. The closest list entry on Sume for a cost view is wan-3.0 at 720p: $0.625 for five seconds. No latency comparison is made here, because the pages cited do not give one for Sume.

Planning rule

Use the table to size a queue, not to promise a delivery time.

  • Batch workloads: one GPU per clip, many clips in flight.
  • Interactive previews: more GPUs per clip, accepting the extra GPU-seconds.
  • Occasional use: a hosted job avoids the idle GPU entirely.
  • License first: the Hunyuan license excludes the EU, UK and South Korea.

What the README does not say

The README gives latency for one resolution, frame count and step count. It gives no figure for 544p, for fewer steps, or for the FP8 weights, so do not extrapolate the table to other settings. It also reports a speedup, not a price, so the cost of a clip depends on your hourly GPU rate, which this post does not assume.

Sources

Related posts

More in Models

All Models posts

Written by Sume