QwenImage21Pipeline in diffusers vs an HTTP image request on Sume

QwenImage21Pipeline runs Qwen-Image-2.1 locally on a GPU. Sume has no 2.1 id: a hosted call is one POST /v1/images with a listed model. Side-by-side.

4 min readSume
All posts

QwenImage21Pipeline is the Diffusers class for running Qwen-Image-2.1 on your own GPU; the README says Diffusers has supported it since day 0, 2026-09-20 (PR 14804). Sume does not list a Qwen-Image-2.1 id, so the hosted equivalent of a pipeline call is a POST /v1/images with a model from the catalog.

Pipeline facts are from the Qwen-Image-2.1 README, read 2026-10-01. Sume request facts are from the Image API, read 2026-10-01.

What does the local call look like?

The README loads the pipeline with QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda"), then calls it with prompt, num_inference_steps=40 and a seeded generator. It lists 2048x2048 for 1:1 and 2752x1536 for 16:9, and says up to 10 reference images work for editing. Other runtimes named in its news list: ComfyUI, vLLM-Omni, SGLang and LightX2V.

Local pipeline call versus Sume image request, README and docs read 2026-10-01.
AspectQwenImage21PipelineSume POST /v1/images
Where it runsYour CUDA GPUHosted job
Model choiceQwen/Qwen-Image-2.1 weightsA catalog id such as qwen/qwen-image-max
Sizewidth and height in pixelsresolution tier and aspect_ratio; explicit pixels not served in v1
Seedmanual_seed(42) on the generatorseed not served in v1
Referencesimage=[...], up to 10 per the READMEinput_references array, range set per model in the catalog

What is the hosted version?

Send model, prompt, optionally aspect_ratio and resolution. Read the model's supported parameters from GET /v1/images/models first, because a parameter a model does not list is rejected with 400 unsupported_parameter.

const res = await fetch("https://api.sume.com/v1/images", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.SUME_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "qwen/qwen-image-max",
    prompt: "A neon shop sign, rainy night, reflections on wet pavement",
    aspect_ratio: "16:9",
  }),
});
console.log(res.status, await res.json());

Which one should I use?

Pick the local pipeline when you need the 2.1 weights themselves; check the Qwen Research License first, since it limits use to research or evaluation without a separate commercial license (see the license post). Pick the hosted request when you want a job over HTTP with the ids Sume lists; Qwen Image 2.1 API covers what is available.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume