QwenImage21Pipeline in diffusers vs an HTTP image request on Sume
QwenImage21Pipeline runs Qwen-Image-2.1 locally on a GPU. Sume has no 2.1 id: a hosted call is one POST /v1/images with a listed model. Side-by-side.

QwenImage21Pipeline is the Diffusers class for running Qwen-Image-2.1 on your own GPU; the README says Diffusers has supported it since day 0, 2026-09-20 (PR 14804). Sume does not list a Qwen-Image-2.1 id, so the hosted equivalent of a pipeline call is a POST /v1/images with a model from the catalog.
Pipeline facts are from the Qwen-Image-2.1 README, read 2026-10-01. Sume request facts are from the Image API, read 2026-10-01.
What does the local call look like?
The README loads the pipeline with QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda"), then calls it with prompt, num_inference_steps=40 and a seeded generator. It lists 2048x2048 for 1:1 and 2752x1536 for 16:9, and says up to 10 reference images work for editing. Other runtimes named in its news list: ComfyUI, vLLM-Omni, SGLang and LightX2V.
| Aspect | QwenImage21Pipeline | Sume POST /v1/images |
|---|---|---|
| Where it runs | Your CUDA GPU | Hosted job |
| Model choice | Qwen/Qwen-Image-2.1 weights | A catalog id such as qwen/qwen-image-max |
| Size | width and height in pixels | resolution tier and aspect_ratio; explicit pixels not served in v1 |
| Seed | manual_seed(42) on the generator | seed not served in v1 |
| References | image=[...], up to 10 per the README | input_references array, range set per model in the catalog |
What is the hosted version?
Send model, prompt, optionally aspect_ratio and resolution. Read the model's supported parameters from GET /v1/images/models first, because a parameter a model does not list is rejected with 400 unsupported_parameter.
const res = await fetch("https://api.sume.com/v1/images", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "qwen/qwen-image-max",
prompt: "A neon shop sign, rainy night, reflections on wet pavement",
aspect_ratio: "16:9",
}),
});
console.log(res.status, await res.json());Which one should I use?
Pick the local pipeline when you need the 2.1 weights themselves; check the Qwen Research License first, since it limits use to research or evaluation without a separate commercial license (see the license post). Pick the hosted request when you want a job over HTTP with the ids Sume lists; Qwen Image 2.1 API covers what is available.
Sources
Related posts
More in Developers
- Qwen Image 2.1 LoRA redistribution: notice file and Hangzhou courts
Sharing Qwen-Image-2.1 weights or a LoRA needs the license copy, modified-file notices and a Notice file; disputes go to the Hangzhou courts. Clause summary.
- remove.bg API rate limit: 500 per minute, weighted by megapixels
remove.bg allows 500 images per minute at about 1 MP, less for larger inputs. Sume limits requests per minute per key and sends retry-after on 429.
- remove.bg bg_color replacement: Sume RMBG gives a transparent PNG
remove.bg's API has bg_color and bg_image_url. Sume RMBG 1.0 takes only an image_url and returns a PNG with alpha, so the new background is a second step.
- remove.bg crop and roi parameters: what Sume RMBG does instead
remove.bg has crop, roi, crop_margin, scale and position. Sume RMBG takes an image_url only and rejects other fields, so cropping happens before or after.
Written by Sume