Prism 2K audio-video needs 4+ 80 GB GPUs; Sume ids at 2K and 4K

Tencent Hunyuan's Prism does native 2K video-audio but wants four or more 80 GB GPUs. Which hosted Sume ids reach 2K or 4K, and what a 10-second clip costs.

5 min readSume
All posts

Tencent Hunyuan's Prism needs four or more 80 GB NVIDIA GPUs to run inference at 1080p or 2K, and one 80 GB GPU for 720p, according to its README read on 2026-10-11. If you would rather call a hosted model, two Sume ids reach the top sizes: minimax-h3 lists 2K and 4K, and gemini-omni-flash-1.1 lists 4K.

Prism is a research release, not a service. The numbers below come from its repository and from Sume's catalog code, and they do not say which output looks better.

What the Prism README states

The Prism repository describes a dynamic sparse attention framework for natively training joint video-audio models at 2K. It supports text-to-video and image-to-video through a reference image path argument. Its README lists project page, code, technical report and a preview model checkpoint under a placeholder date, and the excerpt I read does not name the license type, so check the LICENSE file before any commercial use.

Prism resolutions and inference hardware, from the README (read 2026-10-11)
TierSize listedMax framesInference hardware
720p720 x 1280289single 80 GB GPU
1080p1072 x 19202894 or more 80 GB GPUs
2K1440 x 25602654 or more 80 GB GPUs

Sume ids that list 2K or 4K

From the catalog code, read on 2026-10-11, only two ids that this page can verify list a size above 1080p. Sume's video docs say minimax-h3 renders native 480p/768p and that Sume bills 2K and 4K upscales when a request includes them, so its 2K is not the native 2K Prism describes. The docs list gemini-omni-flash-1.1 at 360p/720p/1080p/4K for 3 to 10 seconds in 16:9 or 9:16.

Hosted ids that list 2K or 4K, with catalog rates (catalog code, read 2026-10-11)
Sume id and sizePer second10 s clipNotes
minimax-h3, 2K$0.1625$1.6255-15 s, native 768p then upscale
minimax-h3, 4K$0.20$2.005-15 s, native 768p then upscale
gemini-omni-flash-1.1, 4K$0.375$3.753-10 s, 16:9 or 9:16, audio always on

Portrait and clip length

Prism's listed sizes are portrait: 720 x 1280, 1072 x 1920 and 1440 x 2560. Both Sume ids accept 9:16, so a vertical feed clip is a request value. Prism's limit is a frame count, not a duration, and the README excerpt does not give a frame rate, so I will not convert 289 frames into seconds. On Sume the limits are in seconds: 15 for minimax-h3 and 10 for the Omni id.

Which route fits

A lab with a multi-GPU node and a need to study or modify a joint audio-video model gets that from Prism. A team that needs a clip tomorrow, billed per second and returned as a URL, gets that from a hosted id. The sizes are similar on paper, but the native-versus-upscale difference above is exactly what to test: render one prompt per candidate and compare the detail at full frame.

If your target is only a vertical clip at 1080p, you do not need either extreme. Several Sume ids list 1080p with audio, such as gemini-omni-flash-1.1 at $0.1875 per second, and a lower tier is often enough for a phone screen.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume