Choose an image model in code: filter GET /v1/images/models
Pick a Sume image model by what the job needs: references, 4:5, a mask, transparency. A Python filter over the catalog descriptors, with the price lookup.

To choose a Sume image model in code, call GET /v1/images/models, keep the rows whose supported_parameters cover what the job needs (reference count, ratio, mask_url, background), then read pricing from the endpoints record of the survivors and take the cheapest. The catalog is the source of truth, so nothing in your code needs to know model ids in advance.
That removes the usual failure where a hard-coded model id meets a parameter it does not list and returns a 400.
What can the catalog filter on?
Each row publishes typed descriptors: enum for a closed list, range for an integer between a min and a max, and boolean for a parameter that is simply present. A parameter that is absent is not supported, and a request that sets it is rejected with 400 unsupported_parameter instead of being dropped silently.
| Job needs | Check in supported_parameters | Rows that pass today |
|---|---|---|
| Edit with 3 or more references | input_references.max >= 3 | All edit-capable rows |
| Instagram portrait | aspect_ratio.values contains 4:5 | GPT Image, Nano Banana, Seedream, FLUX.2, Qwen Image, Ideogram, Recraft V4 |
| Region edit with a mask | mask_url present | The two GPT Image 2.5 ids |
| Transparent background | background present | The two GPT Image 2.5 ids |
| Top quality tier | quality.values contains max | The two GPT Image 2.5 ids |
What does the filter look like?
The script below takes a list of requirements, filters the catalog, then fetches the endpoint record of each remaining model for its output_image price. It uses plain requests, so it runs as is with SUME_API_KEY set. Prices on the endpoint record are already what your wallet is charged, since Sume's margin is applied, and the docs say cost_usd x n is what you pay.
import os, requests
BASE = "https://api.sume.com/v1/images/models"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def get(url):
r = requests.get(url, headers=H, timeout=30)
r.raise_for_status()
return r.json()
def passes(sp, refs, ratio, params):
if sp.get("input_references", {}).get("max", 0) < refs:
return False
if ratio and ratio not in sp.get("aspect_ratio", {}).get("values", []):
return False
return all(p in sp for p in params)
def cheapest(refs=0, ratio=None, params=()):
best = None
for m in get(BASE)["data"]:
if not passes(m["supported_parameters"], refs, ratio, params):
continue
ep = get(f"{BASE}/{m['id']}/endpoints")["endpoints"][0]
cost = ep["pricing"][0]["cost_usd"]
if best is None or cost < best[1]:
best = (m["id"], cost)
return best
print(cheapest(refs=3, ratio="4:5"))What does the cheapest result miss?
Price per image is one line, and for the GPT Image 2.5 ids it is an estimate: the docs say the catalog list is the high-quality 1024 output only, and that admission includes estimated input tokens. A cheapest-model filter will therefore under-price a large quality: max request. Treat the number as a ranking signal, not a quote.
The filter also cannot see quality. A model that passes every descriptor can still render dense text badly. Add a test prompt for your own use case and keep the result, or compare a few rows with the same-prompt comparison script before you fix the choice.
How do I keep it from drifting?
Cache the catalog for a short time and log the chosen id with every job, using the metadata field, which is stored on the job and not sent to the provider. Models are added and retired, so a nightly run of the catalog diff script tells you when your filter would now pick something different.
If you would rather not choose at all, model: "sume/auto" lets Image Router pick, but it never discloses which family ran, and it is not listed in the catalog. Use a filter when you need to know and pin the model, and Auto when you do not.
What is the smallest version I can ship?
Start with two checks: a reference count and a ratio. Those two cover most edit jobs and catch the commonest 400s. Add mask_url and background only when your product offers masked edits or transparent output, since only the two GPT Image 2.5 ids pass those today.
Return the filter result with its reason, such as the descriptor that excluded each model. When a job fails to find a model, the reason tells you whether to relax the ratio, reduce the references, or add a tool outside the Image API.
Sources
Related posts
More in Developers
- Claude Code routines: offset schedules and a Sume Format run
New Claude Code routine schedules start a few minutes past the hour, and unstarted runs show Failed. Key each Sume Format run to the tick so retries are safe.
- Claude structured outputs drop minimum and maxLength; Sume keeps them
Anthropic lists minimum, maxLength and recursion as unsupported in structured outputs. Sume's output_schema accepts the first two; here is what differs.
- Cloud Run concurrency target GA: a Sume webhook receiver sizing
Cloud Run's custom target concurrency and CPU scaling is GA. Size a Sume webhook receiver around the 10 second attempt timeout and 10 retries.
- Cloud Run job delay up to 12 hours: schedule Sume bulk runs
Cloud Run now lets a job start up to 12 hours late (Preview). When that beats a scheduler for Sume bulk runs, and what to keep out of a delayed container.
Written by Sume