Qwen Image vs Qwen Image Max on Sume: $0.025 edit vs $0.094
Qwen Image on Sume costs $0.025 and edits photos with 10 references; Qwen Image Max costs $0.09375 and is text-to-image only. Which to use for what.

Answer
Qwen Image (qwen/qwen-image) costs $0.025 per image on Sume and accepts up to 10 reference images, while Qwen Image Max (qwen/qwen-image-max) costs $0.09375 and is text-to-image only. Max is 3.75 times the price and cannot edit a photo, so you only pay for it when text-only quality is the goal.
Both rows list the same 13 aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 3:2, 2:3, 21:9, 9:21, 1:2 and 2:1) and n up to 4.
| Field | Qwen Image | Qwen Image Max |
|---|---|---|
| Price per image | $0.025 | $0.09375 |
| Reference images | up to 10 | 0, text-only |
| n | 1 to 4 | 1 to 4 |
| Aspect ratios | 13 | 13 |
| Cost of 400 images | $10.00 | $37.50 |
What a reference request does on Max
Max advertises input_references with a range of 0 to 0, and its catalog description says image URLs are not accepted. A request that includes references is rejected with a 400, not silently ignored. That makes it safe to try, because the rejection is not billed.
If you build a model picker, read the descriptor instead of hard-coding names. A model is edit-capable when its input_references max is above 0, and GET /v1/images/models returns that value for every row.
Decision rules
- Edit, background swap or product shot from an existing photo: Qwen Image.
- Pure prompt-to-image with no source photo, where you want to test the top Qwen tier: Qwen Image Max.
- Draft cheaply, finish elsewhere: draft on Qwen Image at $0.025, then rerun the keeper on a model with your preferred look.
- Need text rendering: look at Ideogram 4.5 or ChatGPT Image 2.5 instead; they are the rows that document text and masks.
Check live prices with GET /v1/images/models/qwen/qwen-image/endpoints before a big run. Docs: Sume Image API.
Before a large run
Prices and descriptors change when the catalog changes, so confirm them before you spend. Call GET /v1/images/models/{id}/endpoints for the row you plan to use and read its pricing line and supported_parameters; both come back in one response.
Then run a pilot of three to five images and read usage.cost on each response. Multiply by your planned count for a forecast you can trust. Completed generations are billed in full and failed or cancelled ones are not, so a pilot that errors costs nothing.
For big batches, use mode: "async" or mode: "webhook" with a public HTTPS webhook_url, so no request waits on the 30-second sync limit. Poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result when the job completes.
Sources
Related posts
More in Comparisons
- LiveTranslate 2.3 s latency vs STT plus TTS
Qwen3.8-LiveTranslate cut average lagging from 2.8 to 2.3 seconds. Sume has no live interpreter, but stt_create and tts_create can build an offline dub.
- Recast, Omni edit, Genjutsu or a new shot: which ad variant method?
Four ways to make an ad variant on Sume, limits side by side: Recast swaps people, Omni edit changes by prompt, Genjutsu moves motion, a new shot starts over.
- Recast one holiday ad for four markets: fal vs Sume cost
Four 15-second Recast runs cost $18.00 at fal's $0.30 a second and $22.50 on Sume at $0.375 a second, both at 768p. See the math and what the gap buys.
- Capacity fallbacks: Replicate's model swap vs pinning on Sume
Replicate lists a model falling back to another at capacity. On Sume you pin a model or send sume/auto, and a capacity error is retried with the same key.
Written by Sume