HunyuanVideo 720p: 1,904 s on one GPU, 338 s on eight
README latency for a 1280x720, 129-frame, 50-step clip: 1,904 s on 1 GPU and 338 s on 8. The GPU-seconds each setup burns, and what it means for hosted jobs.

Tencent's HunyuanVideo README reports latency for a 1280x720 clip of 129 frames at 50 steps: 1,904.08 seconds on one GPU, 934.09 on two, 514.08 on four and 337.58 on eight. That is a 5.64x speedup on eight GPUs, but the eight-GPU run burns 2,700.64 GPU-seconds against 1,904.08 for one. Faster wall-clock time costs more total GPU time. Numbers from the README, read 2026-10-08; the GPU-seconds column is multiplication.
The table
The README gives the first three data columns (GPUs, latency, speedup). The last column is latency times the GPU count, which is what you pay for if you rent by the GPU-second.
| GPUs | Latency | Speedup vs 1 GPU | GPU-seconds (latency x GPUs) |
|---|---|---|---|
| 1 | 1904.08 s | 1.00x | 1,904.08 |
| 2 | 934.09 s | 2.04x | 1,868.18 |
| 4 | 514.08 s | 3.70x | 2,056.32 |
| 8 | 337.58 s | 5.64x | 2,700.64 |
How to read it for a render queue
If you need one clip as fast as possible, use eight GPUs: about 5.6 minutes. If you need many clips and throughput matters more than latency, one GPU per clip is the cheapest per clip, because parallel efficiency falls as GPUs are added. By my arithmetic from the README numbers, efficiency (1,904.08 divided by GPU-seconds) is about 102 percent at two GPUs, 93 percent at four and 71 percent at eight.
Memory still applies: the same README says 60 GB is the minimum for 720x1280 at 129 frames, so 'one GPU' means a large-memory GPU.
What a hosted job changes
On a hosted API the batching, GPU count and queue are the provider's problem. Sume's video flow returns a job id at once and you poll GET /v1/videos/{jobId} or use a webhook. Sume's best-practices note says generation usually takes from 30 seconds to several minutes depending on model and parameters, and Sume's admission docs describe queued states when capacity or balance is short.
Sume does not list HunyuanVideo, so this is not a like-for-like swap. The closest list entry on Sume for a cost view is wan-3.0 at 720p: $0.625 for five seconds. No latency comparison is made here, because the pages cited do not give one for Sume.
Planning rule
Use the table to size a queue, not to promise a delivery time.
- Batch workloads: one GPU per clip, many clips in flight.
- Interactive previews: more GPUs per clip, accepting the extra GPU-seconds.
- Occasional use: a hosted job avoids the idle GPU entirely.
- License first: the Hunyuan license excludes the EU, UK and South Korea.
What the README does not say
The README gives latency for one resolution, frame count and step count. It gives no figure for 544p, for fewer steps, or for the FP8 weights, so do not extrapolate the table to other settings. It also reports a speedup, not a price, so the cost of a clip depends on your hourly GPU rate, which this post does not assume.
Sources
Related posts
More in Models
- Hy Image 3.5 lists a 1.5K tier. No Sume image model does
Hy Image 3.5 Preview offers 1K, 1.5K, 2K and 4K. Sume's tier lists are 512, 1K, 2K, 4K or 1K/2K only. For about 1.5K, GPT rows take 1536x1536 as image_size.
- Hy Image 3.5 Preview: 2K or 4K? Tencent and OpenRouter differ
Tencent's launch post says up to 2K; OpenRouter's page lists resolution 1K to 4K. How to test it, and which Sume image rows list a 4K tier or pixel size.
- Hy Image 3.5 lists 13 ratios; GPT Image 2.5 on Sume lacks six
OpenRouter's Hy Image 3.5 Preview ratio list has six ratios GPT Image 2.5 does not offer on Sume: 1:2, 1:4, 2:1, 2:3, 3:2 and 4:1. Which Sume rows list them.
- Hy Image 3.5 Preview: n=1 and seed, against Sume's n and seed
OpenRouter lists Hy Image 3.5 Preview with n capped at 1 and a seed. Sume's Image API takes n up to 4 on most rows and returns 400 for seed. A field table.
Written by Sume