Chinese and English text in images: Hy Image 3.5's claim, Sume's rows
OpenRouter says Hy Image 3.5 Preview renders Chinese and English text. Sume does not list it. How to test Chinese signage on the Sume rows that are listed.

If you need Chinese and English text inside an image, Tencent's Hy Image 3.5 Preview is pitched for it, but Sume does not list that model. On Sume you test the rows that are listed: Nano Banana 2.1 at $0.10 for 1K, ChatGPT Image 2.5, or Imagen 4 Ultra at $0.075.
What the vendor page says
OpenRouter's model page, read 2026-10-08, calls Hy Image 3.5 Preview strong at "rendering Chinese and English text inside images", built on an 80B mixture-of-experts Hy Image 3.0 base, with up to 20 reference images and output up to 4K. Google's pricing page says only that Nano Banana 2.1 has "accurate text rendering" and does not name languages. Neither is a benchmark.
| Model | Vendor-page statement | On Sume |
|---|---|---|
| Hy Image 3.5 Preview | Chinese and English text inside images | Not listed |
| Nano Banana 2.1 | Accurate text rendering, no languages named | google/nano-banana-2.1 |
A fair test on Sume
Write 10 strings you need: shop signs, a price line with yuan symbols, a bilingual menu header. Send each to the listed rows and mark every character error. Keep the prompt text in quotes and say which language each string is in. A 10-string test costs $1.00 on Nano Banana 2.1 at 1K and $0.75 on Imagen 4 Ultra.
If the listed rows fail
If characters come out wrong, render the text outside the model: generate the picture without the text and set the type in your editor or with an overlay in code. That is slower per image but exact. Do not assume a model Sume does not list will be added; check the catalog.
Cost of the fallback test
The full test for one language pair is cheap. Ten strings on Nano Banana 2.1 at 1K cost 10 x $0.10 = $1.00. The same ten on Imagen 4 Ultra cost 10 x $0.075 = $0.75. Add GPT Image 2.5 at xhigh and 1024x1024, about 10 x $0.117 = $1.17 for output only, and the three-model test is under $3.
Score character by character, not by eye. A missing stroke in a Chinese character can change the word, and a reviewer who does not read the language will miss it. Have a native reader check the winners, and keep the failing outputs as a regression set for the next model you test.
Keep the scores with the date, since a model can change under the same name.
What not to conclude
Do not conclude that Hy Image 3.5 is better at Chinese text than Nano Banana 2.1. The only evidence here is two vendor statements. One names Chinese and English, the other names no language. Absence of a claim is not a weakness, and a claim is not a result.
The right conclusion is narrower: if Chinese signage is your core need and the model is not on Sume, you can either call the vendor directly or test the Sume rows and fall back to adding the text in code.
Sources
Related posts
More in Models
- Clef context: 65,536 on Cloudflare, 16,384 on the model card?
Cloudflare's Workers AI page lists 65,536 tokens of context for Clef; its Hugging Face card says 16,384 by default. Budget from the smaller. Cost math inside.
- Do you have to credit an open video model? Label, notice, license
H3 requires a visible 'MiniMax H3' label in commercial products. Hunyuan only encourages 'Powered by'. LTX and Wan ask for notices and a license copy. Per file.
- Does a vertical 9:16 Omni Flash clip cost more than 16:9?
No: on Sume a 9:16 and a 16:9 Gemini Omni Flash 1.1 clip cost the same at each resolution and length. The side-by-side numbers and how to request each ratio.
- Does Sume offer Haiku 5.5, GLM 5.3 Flash or Mistral Large 4?
Checked against Sume's model catalog on 2026-10-08: Haiku 5.5 and GLM 5.3 Flash have enabled rows; Mistral Large 4, Clef and Strands Decider 2B have none.
Written by Sume