Does adding references help? A 3 versus 10 reference test plan on Sume

FLUX 3 Image takes ten references. Before you assume more is better, run 3 against 10 on your own products. A scoring sheet and script, no invented results.

6 min readSume
All posts

The only way to know whether ten references beat three is to test it on your own products with a score sheet, and the sheet below gives you the structure. I have no results to report: the code that follows uses placeholder rows that you replace with your reviewers' scores. Any number printed from them is a demo of the script, not a finding about a model.

The question is live because FLUX 3 Image advertises up to ten input references (read 2026-10-03), and Sume's edit models accept up to ten, with ChatGPT Image 2.5 at 16. More inputs are an opportunity and a cost, since each one adds a thing to reconcile.

Design the test

Pick three to five products with distinct challenges: a label with fine text, a reflective surface, a product with parts. For each, build two reference sets from the same photo pool: three and ten. Hold everything else constant: prompt, model, size, quality. Run each set several times, because one render is an anecdote.

test design, read 2026-10-03
VariableFix or varySetting
ModelFixOne edit-capable Sume row
PromptFixSame text, same legend of image roles
Reference countVary3 and 10
Runs per cellFixAt least 4 per product
Scored fieldsFixLabel ok, shape ok, each 0 or 1
ReviewersFixTwo, scoring blind to the count

Score it blind

Strip file names so reviewers do not know which count produced which image, and have two people score each result with a yes or no on label legibility and on shape fidelity. The script averages those per reference count. The rows are placeholders that illustrate the format; replace them with your data.

import statistics

# placeholder rows: (product, reference_count, reviewer, label_ok 0/1, shape_ok 0/1); replace with your review sheet
REVIEWS = [
    ("mug", 3, "a", 1, 1), ("mug", 3, "b", 0, 1), ("mug", 10, "a", 1, 1), ("mug", 10, "b", 1, 1),
    ("lamp", 3, "a", 0, 1), ("lamp", 3, "b", 0, 0), ("lamp", 10, "a", 1, 1), ("lamp", 10, "b", 1, 0),
]

def pass_rate(count, field):
    rows = [r for r in REVIEWS if r[1] == count]
    return statistics.mean(r[field] for r in rows)

for count in (3, 10):
    print(f"{count:>2} refs  label_ok {pass_rate(count, 3):.0%}  shape_ok {pass_rate(count, 4):.0%}")

Reading the outcome

If ten does not beat three by a margin larger than your reviewer disagreement, the extra references added cost without a gain, and you should use the smaller set. If it wins on label fidelity but loses on shape, split roles: use ten when the label matters and three when the form does. Check each product, because the answer often differs by product type.

Cost matters as much as quality. A request with more references is larger and may take longer, so compare the endpoint pricing lines for your model and note the wall-clock time per run.

Common ways the test misleads

Three traps are worth naming. First, small samples: with two products and two runs each, a single odd render decides the winner. Second, unequal reference quality: the extra seven references in the ten-image set are usually lower quality than the first three, so the test measures reference quality as much as count. Build the larger set by adding good photos, or you are testing the wrong thing. Third, scoring after seeing the count; keep reviewers blind.

Recording the run

Write down the model id, date, size, quality setting and the exact prompt for the test, and keep them with the score sheet. When a vendor ships a new version the same sheet becomes a regression test, and you can see whether the answer to the three versus ten question has changed, or whether a cheaper setting is now good enough.

Where to go next

Keep the sheet and rerun it when a model version changes. For limits per model, see the reference limits comparison and ten references versus Sume's image URL range; for the sixteen-slot case, ChatGPT Image 2.5 references. The parameters are documented at docs.sume.com/models/images.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume