Cheapest Sume image models that take a reference photo
Reference edits start at $0.025 per image on Sume. Ranked list of every catalog model that accepts input_references, with reference limits and prices.

The cheapest reference edits on Sume are Qwen Image and Grok Imagine at $0.025 per image, then Seedream 4 at $0.0325 and Flux 2 Pro at $0.0375 (catalog read 2026-10-03). GPT Image 2.5 takes 16 references, the most of any row, and its medium square estimate is $0.0165, which is cheaper than any of them at that quality. Five rows are text-only: Imagen 4 Fast and Ultra, Recraft V4, Qwen Image Max and Higgsfield Soul accept no references.
Ranked list
The list comes from the catalog descriptors and the per-image endpoint prices (list x 1.25), read 2026-10-03. GPT Image 2.5 prices are repo estimates by quality and size, not billed figures.
| Rank | Model | Price per image | Max references |
|---|---|---|---|
| 1 | Qwen Image | $0.025 | 10 |
| 1 | Grok Imagine | $0.025 | 10 |
| 3 | Seedream 4 | $0.0325 | 10 |
| 4 | Flux 2 Pro | $0.0375 | 10 |
| 4 | Ideogram 4.5, low quality | $0.0375 | 5 |
| 6 | Seedream 5 Lite | $0.04375 | 10 |
| 7 | Seedream 4.5 | $0.05 | 10 |
| 8 | Flux 2 Flex | $0.0625 | 10 |
| 9 | GPT Image 2.5, high (default estimate) | $0.0659 | 16 |
| 10 | Ideogram V3 | $0.075 | 10 |
| 10 | Ideogram 4.5, medium (default) | $0.075 | 5 |
| 12 | Nano Banana 2 | $0.10 | 10 |
| 13 | Nano Banana Pro | $0.1875 | 10 |
Where GPT Image 2.5 fits
GPT Image 2.5 is the only row where quality sets the price. A square at low quality is an estimated $0.0074, medium $0.0165, high $0.0659, xhigh $0.1171 and max $0.2634. It also takes background and mask_url, so it is the row for masked edits and transparent output. The default quality is high, which is why a single catalog figure of $0.0659 shows.
Edit request
Send the photo as input_references and ask for one change.
curl -s https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen/qwen-image", "input_references": [{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}], "prompt": "Keep the person and pose, change the background to a sunny balcony"}' What to test first
Price tells you little about edit fidelity. Run your three hardest pictures through the top three rows and compare the results by eye.
- Does the face or product stay the same?
- Does it follow the one change you asked for?
- How many takes did you need?
Sources
Related posts
More in Pricing
- Opus 5.5 costs 20% less than Opus 5: video agent turn math
Opus 5.5 lists $4 input and $20 output per MTok against Opus 5 at $5 and $25. Priced on a 200K-token video-agent turn, and what Sume did with Opus 5.
- Colossyan Professional $59: cost per minute of NEO and NEO2 video
Colossyan Professional lists $59 a month for 30 NEO minutes and 10 NEO2 minutes: $1.97 per NEO minute, or $1.48 per minute if you use every minute of both.
- How much do 1,000 AI images cost on the Sume image API?
Per-image list prices from the Sume image catalog for 18 model rows, multiplied out to 1,000 images, with the rows where the listed price is only a default.
- Cost of 12 avatar clips (4 each at 15, 30, 45 s) on Sume
Twelve talking-avatar clips, 360 seconds total, cost $66.24 on Standard, $88.20 on Plus and $198.00 on Max at Sume's listed rates, plus $0.95 once.
Written by Sume