Prism 2K audio-video needs 4+ 80 GB GPUs; Sume ids at 2K and 4K
Tencent Hunyuan's Prism does native 2K video-audio but wants four or more 80 GB GPUs. Which hosted Sume ids reach 2K or 4K, and what a 10-second clip costs.

Tencent Hunyuan's Prism needs four or more 80 GB NVIDIA GPUs to run inference at 1080p or 2K, and one 80 GB GPU for 720p, according to its README read on 2026-10-11. If you would rather call a hosted model, two Sume ids reach the top sizes: minimax-h3 lists 2K and 4K, and gemini-omni-flash-1.1 lists 4K.
Prism is a research release, not a service. The numbers below come from its repository and from Sume's catalog code, and they do not say which output looks better.
What the Prism README states
The Prism repository describes a dynamic sparse attention framework for natively training joint video-audio models at 2K. It supports text-to-video and image-to-video through a reference image path argument. Its README lists project page, code, technical report and a preview model checkpoint under a placeholder date, and the excerpt I read does not name the license type, so check the LICENSE file before any commercial use.
| Tier | Size listed | Max frames | Inference hardware |
|---|---|---|---|
| 720p | 720 x 1280 | 289 | single 80 GB GPU |
| 1080p | 1072 x 1920 | 289 | 4 or more 80 GB GPUs |
| 2K | 1440 x 2560 | 265 | 4 or more 80 GB GPUs |
Sume ids that list 2K or 4K
From the catalog code, read on 2026-10-11, only two ids that this page can verify list a size above 1080p. Sume's video docs say minimax-h3 renders native 480p/768p and that Sume bills 2K and 4K upscales when a request includes them, so its 2K is not the native 2K Prism describes. The docs list gemini-omni-flash-1.1 at 360p/720p/1080p/4K for 3 to 10 seconds in 16:9 or 9:16.
| Sume id and size | Per second | 10 s clip | Notes |
|---|---|---|---|
| minimax-h3, 2K | $0.1625 | $1.625 | 5-15 s, native 768p then upscale |
| minimax-h3, 4K | $0.20 | $2.00 | 5-15 s, native 768p then upscale |
| gemini-omni-flash-1.1, 4K | $0.375 | $3.75 | 3-10 s, 16:9 or 9:16, audio always on |
Portrait and clip length
Prism's listed sizes are portrait: 720 x 1280, 1072 x 1920 and 1440 x 2560. Both Sume ids accept 9:16, so a vertical feed clip is a request value. Prism's limit is a frame count, not a duration, and the README excerpt does not give a frame rate, so I will not convert 289 frames into seconds. On Sume the limits are in seconds: 15 for minimax-h3 and 10 for the Omni id.
Which route fits
A lab with a multi-GPU node and a need to study or modify a joint audio-video model gets that from Prism. A team that needs a clip tomorrow, billed per second and returned as a URL, gets that from a hosted id. The sizes are similar on paper, but the native-versus-upscale difference above is exactly what to test: render one prompt per candidate and compare the detail at full frame.
If your target is only a vertical clip at 1080p, you do not need either extreme. Several Sume ids list 1080p with audio, such as gemini-omni-flash-1.1 at $0.1875 per second, and a lower tier is often enough for a phone screen.
Sources
Related posts
More in Comparisons
- Prism 2K video-audio needs 4+ GPUs: the hosted route up to 4K
Tencent Hunyuan's Prism makes 2K video with audio but needs a single 80 GB GPU at 720p and 4+ GPUs above. What Sume lists for 1080p and 4K sound clips instead.
- Runway fails a redirected media URL; Sume follows up to five hops
Runway's API fails a media URL that answers 3XX. Sume follows up to 5 redirects if every hop is public HTTPS. Compare both URL rule sets before you port.
- Same video model, different limits: OpenRouter vs Sume rows
Ten video models are on both OpenRouter and Sume; eight differ in duration or resolution. Veo, Kling, Seedance, Hailuo, Grok side by side, read 2026-10-11.
- Prism audio tags vs Sume generate_audio and audio references
Prism's README describes music, sfx and speech tags. Sume gives a generate_audio flag and audio references on some ids. Which ids, and the limits.
Written by Sume