Compare a local open-weights image model with hosted ones fairly

Same prompt, same shape, no seed: how to test a local Ideogram 4 run against Sume's hosted image models, with a script that prints one image URL per model.

5 min readSume
All posts

Fix four things before you judge any pair of images: the prompt text, the aspect ratio, the number of images, and what you do about the seed. A local open-weights model lets you pin a seed and write a JSON prompt; Sume's hosted catalog takes a plain-text prompt and rejects seed. So a fair test uses plain text on both sides and compares several samples, not one lucky image.

The script below sends one prompt to three hosted catalog models through POST /v1/images, handles both the 200 and the 202 answers, and prints a Sume-hosted URL for each. Run your local model on the same prompt and put the files side by side. The Sume facts come from the Image API docs; the local-model facts come from Ideogram 4's Hugging Face card, both read on 2026-10-03.

What has to match, and what cannot?

Some differences are not removable, so write them down instead of pretending they are not there. The table lists each variable and the honest handling.

Variables in a local-versus-hosted image test, read 2026-10-03
VariableLocal open-weights runSume hosted callHandling
Prompt formatIdeogram 4 was trained on structured JSON captions; plain text works (model card)prompt is a plain stringUse plain text on both sides; run the JSON version as a separate arm
SeedSettable locallyseed is in the schema but advertised by no model, so it returns 400 unsupported_parameterDo not compare single images; generate several per arm
Size256 to 2048 pixels a side, multiples of 16, ratio up to 6:1resolution tiers 512, 1K, 2K, 4K and an aspect_ratioPick 1:1 at about 1K on both
CountAny numbern from 1 to 10, with a per-model ceiling in the catalogRequest the same number per arm
CostYour GPU time, plus a licence if commercialPer-image price on the endpoint recordRecord both before judging

How do you run the hosted side?

Docs say POST /v1/images blocks for up to 30 seconds and returns 200 with the images, but a slow generation returns 202 and a job envelope instead. Check the status code, not the body shape. For a 202, poll the status_url until terminal is true, then read the artifacts from the result_url.

import os, time, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
P = "A bakery window with a hand-painted sign that reads OPEN AT SIX"
for m in ["ideogram/ideogram-v3", "qwen/qwen-image", "black-forest-labs/flux.2-pro"]:
    r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
                      json={"model": m, "prompt": P, "aspect_ratio": "1:1"})
    if r.status_code == 202:
        d = r.json()["data"]
        while True:
            s = requests.get(d["status_url"], headers=H).json()["data"]
            if s["terminal"]:
                break
            time.sleep(5)
        if s["sume_status"] != "completed":
            print(m, "ended as", s["sume_status"])
            continue
        res = requests.get(d["result_url"], headers=H).json()["data"]["result"]
        print(m, res["artifacts"][0]["url"])
    else:
        r.raise_for_status()
        print(m, r.json()["data"][0]["url"])

How should you score the results?

Keep the judgement narrow. A single prompt about a bakery sign tells you about signs, not about product photography or portraits. Run the same loop with prompts from your real work, and change only one thing at a time.

  • Blind the files: rename them to random names before anyone judges.
  • Judge text in the image separately from composition, because text rendering is the claim Ideogram makes for its model.
  • Generate at least four samples per arm, since you cannot fix a seed on Sume.
  • Note failures: a failed hosted generation is not billed, according to the Image API docs, but a failed local run still cost GPU time.
  • Record the price per image on each arm so quality is judged against cost.

Does an open-weights win mean you can ship it?

Not on its own. If your local model is Ideogram 4, its weights are licensed for non-commercial use unless you buy a tier, so a win in the test is a reason to price the licence, not permission to publish. The licence post and the cost break-even post cover those two questions. If the hosted arm wins or ties, the decision is simple: stay on the catalog and skip the GPU.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume