Qwen Image Max at 10 cents or Qwen Image at 3: the edit gap
On Sume, Qwen Image Max is text-to-image only at 10 cents; Qwen Image is 3 cents and accepts reference images. Which to pick, and the 7-cent difference.

On Sume, Qwen Image costs 3 cents per image and accepts reference images for edits, while Qwen Image Max costs 10 cents and is text-to-image only. If your job starts from a photo, the cheaper row is the only Qwen choice; if it starts from words, Max costs 7 cents more per image and you should test whether the output earns it.
The two rows side by side
Both rows list the same 13 aspect ratios, from 1:1 and 16:9 to 21:9, 9:21, 1:2 and 2:1, and a maximum of 4 images per call. The differences are price and input.
| Model | List | Times 1.25 | Billed | References |
|---|---|---|---|---|
| Qwen Image | $0.02 | $0.025 | 3 cents | Accepted |
| Qwen Image Max | $0.075 | $0.09375 | 10 cents | Not accepted (text-to-image only) |
| Gap | $0.055 | $0.06875 | 7 cents | - |
How to decide
The catalog does not say Max is better at any particular task, and neither does this post. What it does tell you is the price and the capability gap, and the way to find out is a paired test: run the same 20 prompts on both. That costs 20 x 3 + 20 x 10 = 260 cents, which is $2.60, and gives you a measured answer for your style.
If your prompts are plain scenes and the cheaper row passes, the saving at volume is large: 1,000 images are $30 on Qwen Image against $100 on Max.
- Edits from a photo: Qwen Image only.
- Pure text-to-image where you want to try the higher tier: Qwen Image Max.
- Both support
nup to 4, so a 4-image grid is 12 cents or 40 cents. - Both are pass-through rows; read the endpoint record for the live price.
What a 'text only' row means in code
The docs say that if the input_references descriptor of a model is {"min": 0, "max": 0}, the model is text-to-image only and rejects references. A request to Max with an image attached therefore fails, and a failed generation is not billed.
So in a pipeline that sometimes has a reference photo and sometimes does not, route by input: with a photo, call Qwen Image; without one, call whichever you prefer. A one-line if on input_references avoids the failed call.
Other cheap edit rows
If Qwen Image does not suit, Grok Imagine (grok-image) is also 3 cents and edits, though it returns one image per call, and FLUX.2 pro and Seedream 4.0 are 4 cents. Pick by ratio list and by how the output looks on your own prompts, not by name.
Running the paired test
Write the 20 prompts down first, with the aspect ratio and any text requirements. Send each to both rows with n: 1, name the files by row, and judge them blind by shuffling the filenames. Add up usage.cost; it should come to $2.60. If the cheaper row wins even a third of the time, route that share of traffic to it and keep Max only for the prompts where it clearly helps.
Sources
Related posts
More in Models
- Recraft V4 on Sume returns WebP only: a PNG conversion step, 5 cents
Recraft V4 on Sume bills 5 cents per image, is text-to-image only and lists webp as its only output format. A short Pillow script converts results to PNG.
- Grep for nano-banana-2 in your code: those calls bill as 2.1
Retired Nano Banana 2 ids still run on Sume, as 2.1 at 10 cents at 1K. A grep-and-replace plan so logs, tests and cost sheets name the model that ran.
- Same Omni video twice: no seed, but a replay returns it
Omni Flash lists no temperature or seed, and Sume rejects seed on every video model. To repeat a clip, replay an idempotency key or use a video_url edit.
- Seedance 2.5 costs 1.53x Seedance 2.0 at 720p for 10 seconds
A 10 s 9:16 clip at 720p is $5.78 on Seedance 2.5 and $3.78 on Seedance 2.0 on Sume. What the extra $2.00 buys, and the ratio at 480p and 1080p.
Written by Sume