Qwen-Image 2.1 is 7B: self-host it or call hosted qwen-image on Sume
Qwen-Image 2.1 is a 7B model under the Qwen Research License. Sume hosts qwen-image and qwen-image-max, not 2.1; hosted costs $0.025 or $0.094 per image.

Qwen-Image 2.1 is a 7-billion-parameter visual generator, and its Hugging Face model card (read 2026-10-05) names the Qwen Research License Agreement. Sume does not host 2.1. Sume's catalog has qwen/qwen-image and qwen/qwen-image-max, which are other rows, listed at $0.02 and $0.075 per image before margin. Read the licence before you plan any commercial use of the weights.
What the model card says
Facts from the card, read 2026-10-05:
| Item | Value |
|---|---|
| Size | 7B visual generator |
| Tasks | Text-to-image and editing in one model |
| Transparency | Native RGBA |
| Reference images | Up to 10 |
| Max size | Up to 2048 x 2048 |
| Ratios | 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16 |
| Inference steps | 40 |
| Library | Diffusers compatible |
| Licence | Qwen Research License Agreement |
Hosted cost on Sume
Sume bills provider list x 1.25. These are the Sume rows, not Qwen-Image 2.1, and the arithmetic is before cent rounding. The live endpoints pricing line is authoritative.
| Sume model id | List | x 1.25 | 1,000 images |
|---|---|---|---|
| qwen/qwen-image | $0.02 | $0.025 | $25 |
| qwen/qwen-image-max | $0.075 | $0.09375 | $93.75 |
Self-host or call an API
Self-hosting means you provide a GPU, handle the Diffusers install, and carry the licence question. We did not measure memory use or speed, and the model card text we read does not give a verified VRAM figure, so test on your own hardware before you size anything. A hosted call costs by the image and needs nothing local, but you get the hosted model, not 2.1.
A rough break-even: if the GPU you would rent costs $X per hour, divide by your images per hour to get a per-image figure, then compare it with $0.025. Include idle time, which is where self-hosting usually loses.
A hosted call
This calls the hosted row on Sume. Because 2.1 is not on Sume, results will differ from the 2.1 weights.
import asyncio, os, httpx
async def main():
body = {"model": "qwen/qwen-image", "prompt": "a bakery storefront at sunrise, poster style", "aspect_ratio": "3:4"}
headers = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
async with httpx.AsyncClient(timeout=60) as c:
r = await c.post("https://api.sume.com/v1/images", headers=headers, json=body)
print(r.status_code, r.json().get("usage"))
asyncio.run(main())What we did not check
We did not run the model, so this post makes no claims about speed, memory or output quality. Hosted and self-hosted are two different products here: the hosted rows are not Qwen-Image 2.1, so a hosted result tells you nothing certain about 2.1.
Sources
Related posts
More in Models
- Recraft V4.1 from $0.035 vs V4 from $0.04: Sume lists only V4
Recraft's docs list V4.1 from $0.035 per image and V4 vectors from $0.04. Sume's catalog has recraft/recraft-v4 only, about $0.05 after its x 1.25 margin.
- 'reference_audio_urls requires a reference image or video'
Audio cannot be the only reference on Sume video models. Add one image or video reference. Which models take audio references, and their caps.
- 'Reference inputs must not exceed 12 total files': per-type caps
Most Sume video models take 9 reference images, 3 videos, 3 audios, 12 files total. Wan 3.0 takes 10, 5 and 5; Gemini Omni takes 10 images and 3 short videos.
- Seedance 2.0, Fast, Mini and 2.5: a 10-second 720p price ladder
Four Seedance ids are callable on Sume. A 10-second 720p clip is about $1.89 on Mini, $3.03 on Fast, $3.78 on 2.0 and $5.78 on 2.5. Limits for each.
Written by Sume