Text-to-image-only models on Sume reject input_references
Five Sume image ids take no reference images: Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max. The error you get and a Python guard before sending.

Not every image model on a router does image-to-image. In the Sume catalog, a model that cannot edit reports input_references as a range from 0 to 0. If you send a reference anyway, the API answers 400 unsupported_parameter with the message that the model is text-to-image only, rather than ignoring the image and spending money on a text-only render.
From the catalog data in the repo, five ids are text-only: higgsfield/soul, google/imagen-4-fast, google/imagen-4-ultra, recraft/recraft-v4 and qwen/qwen-image-max.
Text-only versus edit-capable
| Sume id | input_references max |
|---|---|
| higgsfield/soul | 0 |
| google/imagen-4-fast | 0 |
| google/imagen-4-ultra | 0 |
| recraft/recraft-v4 | 0 |
| qwen/qwen-image-max | 0 |
| qwen/qwen-image | 10 |
A guard
Check the descriptor, then fall back to an edit-capable id if you have references. This keeps a pipeline from failing on a model swap:
import os, requests
headers = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def accepts_refs(model_id: str) -> bool:
url = f"https://api.sume.com/v1/images/models/{model_id}/endpoints"
r = requests.get(url, headers=headers, timeout=30)
r.raise_for_status()
p = r.json()["endpoints"][0]["supported_parameters"]
return p.get("input_references", {}).get("max", 0) > 0
print(accepts_refs("google/imagen-4-ultra"))
print(accepts_refs("openai/gpt-image-2.5"))
Why this matters for trend models
New image releases often split into a text model and an edit model. The same split can appear in a catalog, so treat edit support as a per-id fact and read it from the descriptor every time rather than from a model's name.
How this was checked
Vendor facts come from the pages listed in the sources, read on 2026-10-05. Sume facts come from the Image API docs and the catalog code on main on the same date. Catalogs and limits change, so read the descriptors from GET /v1/images/models before you pin a number in production code.
Sources
Related posts
More in Developers
- Text to speech API with emotion: audition four values, keep the winner
Sume TTS 1.0 takes a free-text emotion up to 64 characters in generation_config. Run a four-take audition for 4 cents and store the winner.
- Text-to-video API in Node: submit, poll and download with a deadline
A runnable Node 18+ script that submits a text-to-video job to Sume, polls it with a deadline, and saves the MP4. No dependencies and no top-level await.
- Threads image post: 8 MB, 320 to 1440 px wide, from video frames
Threads images must be JPEG or PNG, up to 8 MB, 320 to 1440 px wide. Pull a still from a Sume clip with video-frames and clamp the long edge to 1440.
- Threads video aspect ratio up to 10:1 and 1920 px: timeline output
Threads allows video ratios from 0.01:1 to 10:1 and 1920 px, and recommends 9:16. The Sume timeline default 1080x1920 fits, and width and height are settable.
Written by Sume