Why input_references fails on Imagen, Recraft, Soul and Qwen Max
Five Sume image rows list zero input_references: imagen-4 fast and ultra, recraft-v4, Soul and qwen-image-max. Check the catalog in Python first.

A reference-image request fails on imagen-4-fast, imagen-4-ultra, recraft-v4, higgsfield-soul and qwen-image-max because the catalog lists their input_references descriptor as {min: 0, max: 0}: they are text-to-image only. The Image API says such a model rejects references, and any parameter a model does not list returns 400 unsupported_parameter; Sume does not silently drop it.
Check before you send
Read the descriptor once and cache it. This prints each model with its reference ceiling.
import os, requests
r = requests.get(
"https://api.sume.com/v1/images/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
for m in r.json()["data"]:
refs = m["supported_parameters"].get("input_references", {})
print(m["id"], refs.get("max", 0))Reference ceilings and prices
| Model | Reference ceiling | Per image |
|---|---|---|
| openai/gpt-image-2.5 | 16 | $0.0074 |
| google/nano-banana-pro | 10 | $0.1875 |
| bytedance-seed/seedream-5-lite | 10 | $0.0437 |
| black-forest-labs/flux.2-pro | 10 | $0.0375 |
| ideogram/ideogram-v4.5 | 5 | $0.0375 |
| recraft/recraft-v4 | 0 | $0.05 |
| google/imagen-4-fast | 0 | $0.025 |
| qwen/qwen-image-max | 0 | $0.0938 |
| higgsfield/soul | 0 | $0.005 |
What to do instead
For image-to-image, pick a row with a non-zero ceiling. If you want the look of a text-only model applied to your photo, generate with that model and then use a reference-capable row to merge the two.
A defensive wrapper
Fail before you send, not after.
def can_reference(models: list[dict], model_id: str) -> bool:
for m in models:
if m["id"] == model_id:
refs = m["supported_parameters"].get("input_references", {})
return refs.get("max", 0) > 0
return False
print(can_reference([{"id": "a", "supported_parameters": {"input_references": {"max": 0}}}], "a"))The other limits are separate
The reference ceiling is its own descriptor, not the output count n. n is capped at 4 on most rows (1 on grok-image), so a gpt-image-2.5 edit can carry up to 16 references and still return up to 4 images. Rows such as ideogram-v4.5 stop at 5 references, and most other edit-capable rows (nano-banana-2.1, seedream-5-lite, flux-2-pro, qwen-image) allow 10. Read the live descriptor rather than hard-coding any of these.
Sources
Related posts
More in Developers
- Is Higgsfield Genjutsu in your Sume catalog? Check, then price it
higgsfield-genjutsu appears in the catalog only when its provider is configured. List your models, then price a 4 to 30 second source at 480p or 720p.
- Jupyter: move Sora cells to Sume, where a cell rerun is a retry
Re-running a notebook cell resubmits the request. Hold one Idempotency-Key per take in a variable so Sume returns the first video job, and show the file inline.
- Kotlin: submit and poll a Sume video job with java.net.http
A Kotlin port of a Sora videos call: POST /v1/videos with an Idempotency-Key, poll every 30 seconds until done, with the JDK client and kotlinx.serialization.
- Kubernetes CronJob that submits a nightly 30-second Wan 3.0 clip
A CronJob manifest using curlimages/curl and a Secret: one dated Idempotency-Key per night, concurrency forbidden, and a month of reserves at each resolution.
Written by Sume