Compare a local open-weights image model with hosted ones fairly
Same prompt, same shape, no seed: how to test a local Ideogram 4 run against Sume's hosted image models, with a script that prints one image URL per model.

Fix four things before you judge any pair of images: the prompt text, the aspect ratio, the number of images, and what you do about the seed. A local open-weights model lets you pin a seed and write a JSON prompt; Sume's hosted catalog takes a plain-text prompt and rejects seed. So a fair test uses plain text on both sides and compares several samples, not one lucky image.
The script below sends one prompt to three hosted catalog models through POST /v1/images, handles both the 200 and the 202 answers, and prints a Sume-hosted URL for each. Run your local model on the same prompt and put the files side by side. The Sume facts come from the Image API docs; the local-model facts come from Ideogram 4's Hugging Face card, both read on 2026-10-03.
What has to match, and what cannot?
Some differences are not removable, so write them down instead of pretending they are not there. The table lists each variable and the honest handling.
| Variable | Local open-weights run | Sume hosted call | Handling |
|---|---|---|---|
| Prompt format | Ideogram 4 was trained on structured JSON captions; plain text works (model card) | prompt is a plain string | Use plain text on both sides; run the JSON version as a separate arm |
| Seed | Settable locally | seed is in the schema but advertised by no model, so it returns 400 unsupported_parameter | Do not compare single images; generate several per arm |
| Size | 256 to 2048 pixels a side, multiples of 16, ratio up to 6:1 | resolution tiers 512, 1K, 2K, 4K and an aspect_ratio | Pick 1:1 at about 1K on both |
| Count | Any number | n from 1 to 10, with a per-model ceiling in the catalog | Request the same number per arm |
| Cost | Your GPU time, plus a licence if commercial | Per-image price on the endpoint record | Record both before judging |
How do you run the hosted side?
Docs say POST /v1/images blocks for up to 30 seconds and returns 200 with the images, but a slow generation returns 202 and a job envelope instead. Check the status code, not the body shape. For a 202, poll the status_url until terminal is true, then read the artifacts from the result_url.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
P = "A bakery window with a hand-painted sign that reads OPEN AT SIX"
for m in ["ideogram/ideogram-v3", "qwen/qwen-image", "black-forest-labs/flux.2-pro"]:
r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
json={"model": m, "prompt": P, "aspect_ratio": "1:1"})
if r.status_code == 202:
d = r.json()["data"]
while True:
s = requests.get(d["status_url"], headers=H).json()["data"]
if s["terminal"]:
break
time.sleep(5)
if s["sume_status"] != "completed":
print(m, "ended as", s["sume_status"])
continue
res = requests.get(d["result_url"], headers=H).json()["data"]["result"]
print(m, res["artifacts"][0]["url"])
else:
r.raise_for_status()
print(m, r.json()["data"][0]["url"])How should you score the results?
Keep the judgement narrow. A single prompt about a bakery sign tells you about signs, not about product photography or portraits. Run the same loop with prompts from your real work, and change only one thing at a time.
- Blind the files: rename them to random names before anyone judges.
- Judge text in the image separately from composition, because text rendering is the claim Ideogram makes for its model.
- Generate at least four samples per arm, since you cannot fix a seed on Sume.
- Note failures: a failed hosted generation is not billed, according to the Image API docs, but a failed local run still cost GPU time.
- Record the price per image on each arm so quality is judged against cost.
Does an open-weights win mean you can ship it?
Not on its own. If your local model is Ideogram 4, its weights are licensed for non-commercial use unless you buy a tier, so a win in the test is a reason to price the licence, not permission to publish. The licence post and the cost break-even post cover those two questions. If the hosted arm wins or ties, the decision is simple: stay on the catalog and skip the GPU.
Sources
Related posts
More in Comparisons
- Face and body swap video AI: Recast vs Sume's avatar face swap
Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.
- Face swap vs H3 Max Recast: which swaps the person in a video?
Sume offers two ways to put a different person in a video: Avatar Face Swap (Beta) and H3 Max Recast. Inputs, length limits, audio and price side by side.
- FLUX 3 Image alternatives on Sume, feature by feature
No FLUX 3 Image in Sume's catalog yet. Match each FLUX 3 feature, 10 references, 4K, region edits, grounding, to the Sume model that has it or the gap.
- FLUX 3 Image vs Nano Banana Pro for 4K editing on Sume
FLUX 3 Image is not in Sume's catalog; Nano Banana Pro is, with a 4K tier and 10 reference images. A checklist of what each does for editing, from vendor pages.
Written by Sume