Hy Image 3.0 or 3.5 Preview: 3 vs 20 references, size rules, seed
Tencent's Hy Image 3.0 and 3.5 Preview side by side: reference counts, prompt limits, size ranges, seed ranges and resize_max_pixels, with what Sume lists.

Hy Image 3.0 takes up to 3 reference images and an 8,192-character prompt; Hy Image 3.5 Preview takes up to 20 references and a 100k-token limit through a messages protocol, with a much wider size range. Hy Image 3.5 Preview is not listed in the Sume image catalog (read 2026-10-09). Pick 3.5 for edits and larger output, and 3.0 for simple text-to-image with automatic prompt rewriting.
How do the two models differ?
Everything in this table is from Tencent's page.
| Item | Hy Image 3.0 (hy-image-v3) | Hy Image 3.5 Preview (hy-image-v3.5-preview) |
|---|---|---|
| Request | prompt string, v3-generation | messages, v35-generation |
| Prompt limit | 8192 characters | 100k tokens |
| References | 0 to 3, up to 10 MB each | up to 20, up to 20 MB each |
| size | 512 to 2048, area at most 1024x1024 | 256 to 8192, area at most 16777216 |
| seed | 1 to 4294967295, n=1 only | 0 to 2^63-1 |
| Prompt rewrite | revise, revised_prompt returned | no revise field documented |
What is resize_max_pixels?
It is a 3.5 field for the reference side: an input image whose area is at or under the threshold passes as is; a larger one is scaled to roughly the threshold before the model sees it, so a 6K photo does not blow up the context. The default is 1,048,576 pixels. generate_max_pixels is the output side, with 1K, 1.5K (default) and 2K levels, and only applies when you do not pass size.
What can I call on Sume instead?
Neither Hy model is listed. For edits with many references, GPT Image 2.5 takes up to 16 and Nano Banana 2.1 up to 10. Sume does not serve seed, so a fixed-seed reproduction is not possible there; log your prompt and keep the file instead. Rows and limits are in the image models docs.
- Hy 3.5 for multi-turn edits and 4K via
size. - Hy 3.0 for simple text-to-image with automatic prompt rewriting (about 11 seconds).
- Sume: pick a listed row by reference count and price.
Where is the price?
Tencent's page gives no per-image price, so this post states none. Check your Tencent console before you commit.
The 3.5 size range of 256 to 8192 with an area cap is the biggest practical change from 3.0, which stops at 2048 and the area of a 1024 square. If your use case is a poster or a banner, the size rules alone may decide the model. If it is a square avatar, either model is enough, and price decides.
Sources
Related posts
More in Comparisons
- Keyterm prompting costs $0.05/hour at two vendors: 60 hours is $3.00
ElevenLabs and AssemblyAI both list keyterm prompting at $0.05 an hour: $3.00 for 60 hours of calls. Sume STT has a language hint but no keyterm list.
- Kling 3 Pro text-to-video is $0.14/s on fal; four Sume ids cost less
Sume lists no Kling text-to-video, only Kling 3.0 Motion Control at $0.1575. Four Sume text-to-video ids cost $0.075 to $0.125 per second against fal's $0.14.
- LTX-2.5 vs MiniMax H3: license lines and run requirements
LTX-2.5 vs MiniMax H3 from vendor pages: 10M vs 20M USD revenue lines, excluded territories, Python and CUDA needs, frame rules, audio, and what Sume lists.
- Luma Ray 3.2 alternative: one /v1/generations route vs Sume's two
Luma's API sends images and video through POST /v1/generations. Sume splits them into /v1/videos and /v1/images and does not list ray-3.2 in its docs.
Written by Sume