Ten reference images for one product still: what goes in each slot
FLUX 3 Image takes up to ten references, GPT Image 2.5 on Sume takes sixteen, Ideogram 4.5 five. A slot plan and a script that trims it to each model's cap.

For one product still, spend references in priority order: the hero angle first, a flat shot of the label second, then extra angles, a scale object, a lighting reference and a surface. Ten slots is the ceiling FLUX 3 Image advertises, and most edit models on Sume accept up to ten, so a plan that ranks references lets you cut the tail when a model allows fewer.
The limits differ by model, and they are the reason to rank. Per Sume's Image API docs, ChatGPT Image 2.5 takes up to 16 references, other edit models take up to 10, Ideogram 4.5 takes five in total (the first is the image being edited), and a few rows, including Imagen 4 and Recraft V4, are text-to-image only and reject references.
The slot plan
More references do not automatically mean a better result: every slot you fill is something the model has to reconcile. Give each one a job you can name, and tell the model which is which in the prompt. Sume's docs describe references as a list, so "Image 1 is the hero angle" is the way to bind a role to a position.
| Slot | Role | Why it earns the slot | Cut order |
|---|---|---|---|
| 1 | Hero angle | Sets shape and proportions | Never |
| 2 | Label flat | Keeps text and logo legible | Never |
| 3 | Side or back angle | Fixes the parts the hero hides | Late |
| 4 | Scale object (hand, coin) | Stops the product drifting in size | Middle |
| 5 | Lighting reference | Matches direction and softness | Middle |
| 6 | Surface or backdrop | Material of the table or wall | Early |
| 7-10 | Props, mood, brand colors | Style only | First |
Trim the plan to the model
Sume publishes each model's limit as a capability descriptor: GET /v1/images/models returns supported_parameters.input_references with a max. The script reads that cap, keeps the top-ranked slots, builds the reference list, and writes the "Image N is..." legend that goes into the prompt. It reports what it dropped so a smaller cap is never silent.
import os
import requests
def reference_cap(model_id):
r = requests.get(
"https://api.sume.com/v1/images/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
if m["id"] == model_id:
return m["supported_parameters"].get("input_references", {}).get("max", 0)
raise KeyError(model_id)
# highest priority first; the list is cut to whatever the model accepts
SLOTS = [
("hero angle", "https://example.com/sku/front.jpg"),
("label flat", "https://example.com/sku/label-flat.png"),
("side angle", "https://example.com/sku/side.jpg"),
("scale object", "https://example.com/props/hand.jpg"),
("light reference", "https://example.com/looks/window-light.jpg"),
("surface", "https://example.com/looks/slate.jpg"),
]
def build_references(model_id, slots=SLOTS):
cap = reference_cap(model_id)
kept = slots[:cap]
refs = [{"type": "image_url", "image_url": {"url": url}} for _, url in kept]
legend = " ".join(f"Image {i + 1} is the {name}." for i, (name, _) in enumerate(kept))
return refs, legend, [n for n, _ in slots[cap:]]
if __name__ == "__main__":
refs, legend, dropped = build_references("openai/gpt-image-2.5")
print(legend, "| dropped:", dropped)Where this stops working
On an edit model that reads the first image as the source to edit, order matters more than count. Ideogram 4.5 on Sume edits the first reference and uses up to four more as style or content references, so the hero angle is also the canvas. If your plan puts the label flat first, you will edit the label sheet. Check the order for each model before a batch.
A practical habit is to shoot the label flat on a plain background under even light, because a label photographed on a curved bottle bends the text and the model will faithfully copy the bend. If the product has more than one colorway, give each its own run rather than mixing colorways in the reference list, since two similar products in one list invite a blend.
Public HTTPS URLs only. Sume rejects localhost, private addresses and non-HTTPS URLs before submission, so a local file has to be hosted first.
Test it before you trust it
Run the same product with three slots, six and ten on the same prompt, and score label legibility and shape separately. If ten does not beat six on your product, the extra four were cost, not quality. Sume's GPT Image 2.5 reference guide covers the sixteen-slot case in more detail. Keep the three runs, their prompts and the scores in one spreadsheet so the decision is evidence rather than a feeling, and rerun it when a model version changes, since a limit that was a ceiling last month can move.
The model limits are in the Image API docs; the older Image 1.0 page lists its own 1 to 10 range on the retiring alias.
Sources
Related posts
More in Use cases
- Avatar reaction 3.49 vs 3.06 predicted: run your own two-clip test
HeyGen's survey says avatar users saw warmer reactions than skeptics predicted. Test your own audience with two Sume avatar clips, one script and safe retries.
- Test three Reel hooks from one clip: three video trims plus renders
Cut three first-3-second hooks with video trim at $0.02 each, put each ahead of the same body in Timeline, and the test costs $0.36 for 60-second Reels.
- Text-only Shorts on YouTube: when scrolling text counts as value
YouTube lists scrolling text with minimal narrative as not allowed. How to write a text Short with real commentary and time its cues with Sume captions.
- Thanksgiving 2026 takeout video: a preorder cutoff clip
Thanksgiving is Thursday, Nov 26, 2026. Make three short vertical clips for takeout preorders, with the cutoff burned in as caption text, from your own photos.
Written by Sume