Does adding references help? A 3 versus 10 reference test plan on Sume
FLUX 3 Image takes ten references. Before you assume more is better, run 3 against 10 on your own products. A scoring sheet and script, no invented results.

The only way to know whether ten references beat three is to test it on your own products with a score sheet, and the sheet below gives you the structure. I have no results to report: the code that follows uses placeholder rows that you replace with your reviewers' scores. Any number printed from them is a demo of the script, not a finding about a model.
The question is live because FLUX 3 Image advertises up to ten input references (read 2026-10-03), and Sume's edit models accept up to ten, with ChatGPT Image 2.5 at 16. More inputs are an opportunity and a cost, since each one adds a thing to reconcile.
Design the test
Pick three to five products with distinct challenges: a label with fine text, a reflective surface, a product with parts. For each, build two reference sets from the same photo pool: three and ten. Hold everything else constant: prompt, model, size, quality. Run each set several times, because one render is an anecdote.
| Variable | Fix or vary | Setting |
|---|---|---|
| Model | Fix | One edit-capable Sume row |
| Prompt | Fix | Same text, same legend of image roles |
| Reference count | Vary | 3 and 10 |
| Runs per cell | Fix | At least 4 per product |
| Scored fields | Fix | Label ok, shape ok, each 0 or 1 |
| Reviewers | Fix | Two, scoring blind to the count |
Score it blind
Strip file names so reviewers do not know which count produced which image, and have two people score each result with a yes or no on label legibility and on shape fidelity. The script averages those per reference count. The rows are placeholders that illustrate the format; replace them with your data.
import statistics
# placeholder rows: (product, reference_count, reviewer, label_ok 0/1, shape_ok 0/1); replace with your review sheet
REVIEWS = [
("mug", 3, "a", 1, 1), ("mug", 3, "b", 0, 1), ("mug", 10, "a", 1, 1), ("mug", 10, "b", 1, 1),
("lamp", 3, "a", 0, 1), ("lamp", 3, "b", 0, 0), ("lamp", 10, "a", 1, 1), ("lamp", 10, "b", 1, 0),
]
def pass_rate(count, field):
rows = [r for r in REVIEWS if r[1] == count]
return statistics.mean(r[field] for r in rows)
for count in (3, 10):
print(f"{count:>2} refs label_ok {pass_rate(count, 3):.0%} shape_ok {pass_rate(count, 4):.0%}")Reading the outcome
If ten does not beat three by a margin larger than your reviewer disagreement, the extra references added cost without a gain, and you should use the smaller set. If it wins on label fidelity but loses on shape, split roles: use ten when the label matters and three when the form does. Check each product, because the answer often differs by product type.
Cost matters as much as quality. A request with more references is larger and may take longer, so compare the endpoint pricing lines for your model and note the wall-clock time per run.
Common ways the test misleads
Three traps are worth naming. First, small samples: with two products and two runs each, a single odd render decides the winner. Second, unequal reference quality: the extra seven references in the ten-image set are usually lower quality than the first three, so the test measures reference quality as much as count. Build the larger set by adding good photos, or you are testing the wrong thing. Third, scoring after seeing the count; keep reviewers blind.
Recording the run
Write down the model id, date, size, quality setting and the exact prompt for the test, and keep them with the score sheet. When a vendor ships a new version the same sheet becomes a regression test, and you can see whether the answer to the three versus ten question has changed, or whether a cheaper setting is now good enough.
Where to go next
Keep the sheet and rerun it when a model version changes. For limits per model, see the reference limits comparison and ten references versus Sume's image URL range; for the sixteen-slot case, ChatGPT Image 2.5 references. The parameters are documented at docs.sume.com/models/images.
Sources
Related posts
More in Comparisons
- Does Sume have Veo 3.1? No. Which video models it lists instead
Sume's video catalog has no Veo model. See the models it does list, with duration and resolution ranges, and how Google's own docs now steer to Omni Flash.
- Editframe cloud render $0.02 a minute plus $99 vs Sume Timeline
Editframe bills cloud render minutes by megapixel band from $0.02, on a $99 plan. A 1080x1920 Reel is 2.07 MP. Sume renders a minute for $0.10 with no base fee.
- Edits has 250+ fonts: Sume's 29 caption fonts are Hangul-only
Edits ships 250+ fonts and 50+ text animations. Sume Video Captions lists 29 fonts, all Hangul, and no font field for Latin styles. What that means.
- Edits has 30+ caption styles; Sume has 10 plus design overrides
Instagram says Edits offers 30+ caption styles in multiple languages. Sume Video Captions documents 10 named styles and a design override object. See the table.
Written by Sume