Qwen Image vs Qwen Image Max on Sume: edit support and a 3.75x price
Qwen Image bills $0.025 and edits photos; Qwen Image Max bills $0.09375 and is text-to-image only on Sume. Both list 13 ratios, including 21:9 and 9:21.

Qwen Image Max is the dearer row and the one that cannot edit. On Sume, Qwen Image bills $0.025 per image and accepts reference images, while Qwen Image Max bills $0.09375 (3.75 times more) and is text-to-image only. Both list 13 aspect ratios, including 21:9, 9:21, 2:1 and 1:2.
Side by side
| Field | Qwen Image | Qwen Image Max |
|---|---|---|
| API id | qwen/qwen-image | qwen/qwen-image-max |
| List per image | $0.02 | $0.075 |
| Billed (list x 1.25) | $0.025 | $0.09375 |
| Reference images (edit) | Yes | No |
| Aspect ratios | 13 | 13 |
| Custom pixels (image_size) | Yes | Yes |
| Max n per call | 4 | 4 |
Choosing by job
- Editing a photo, restyling a reference, or a first-frame fix: Qwen Image, because Max rejects references.
- A pure prompt-to-image job where you want to test the top Qwen row: Max, at 3.75 times the price.
- A wide banner or a tall strip: either, since both list 21:9, 9:21, 2:1 and 1:2.
What the price gap buys in a batch
Four takes of one prompt (n: 4) cost $0.10 on Qwen Image and $0.375 on Max. A 100-image set is $2.50 against $9.375. These are plain multiples of the billed rates. Whether Max images are worth the gap depends on your prompts, so run a small A/B first; Sume does not publish a quality comparison here.
If your A/B shows Max is better only on a few prompts, route those prompts to Max and the rest to Qwen Image. Pin both ids in your code and choose per request.
Wide and tall work
Qwen Image costs $0.025 and Qwen Image Max $0.09375, a ratio of 3.75. Ask for the ratio you need through aspect_ratio where the row lists it, since cropping a square throws away pixels you paid for.
If you need to start from a photo, only Qwen Image can take it as a reference. For a result from text alone, test both and keep Max for the prompts where Qwen Image misses. Remember that the sync wait on /v1/images is capped at 30 seconds; if a slow call returns 202, poll the job rather than resubmit.
The error you will see
Sending input_references to Qwen Image Max returns 400 unsupported_parameter. The model's input_references descriptor is {min: 0, max: 0}. Read it from GET /v1/images/models and gate the field in your client. The Image API page explains how the descriptors drive validation, and the same rule applies to mask_url, which only ChatGPT Image 2.5 accepts.
Sources
Related posts
More in Models
- Recurring cast in Gemini Omni Flash 1.1: 10 images, 3 short clips
Omni Flash 1.1 reference-to-video on Sume takes up to 10 images plus 3 clips of 3 s or less, addressed as IMAGE_REF_0. A 10 s 9:16 shot at 720p is $1.25.
- Remove an object from a photo without a mask: Ideogram 4.5 prompt edit
Remove a bin, cable or stray person from a photo with no mask: send it to ideogram/ideogram-v4.5 and name the object. $0.0375 to $0.275 per image on Sume.
- Replace one prop in a video with AI: bottle to apple edit prompt
Swap one object in a finished clip with gemini-omni-flash-1.1 on Sume: the docs example prompt, fields you cannot set, and an 8-second price.
- Seedance 2.0 mini, fast, standard vs 2.5: price of a 15-second clip
A 15-second 720p clip on Sume costs $2.84 on Seedance 2.0 mini, $4.54 on fast, $5.67 on standard and $8.67 on 2.5, before the 5.5% fee. Arithmetic shown.
Written by Sume