31 lookbook edits, 3 reference images each: 93 inputs, flat rows
31 edits at 2K with three references each cost $4.65 on Nano Banana 2.1 and $5.81 on Pro: Sume's rows are per image, not per input.

The short answer
Thirty-one lookbook edits at 2K, each sent with three reference images (a garment, a model photo, a background), cost $4.65 on Nano Banana 2.1 and $5.81 on Nano Banana Pro. That is 93 reference inputs in total, and they do not appear in the arithmetic: the catalog row is per output image.
GPT Image 2.5 is the case to watch. Its upstream rates are per token, so confirm what a reference-heavy call costs before you scale it.
How to send the references
On POST /v1/images, put the references in input_references as image_url objects with public HTTPS URLs. Sume rejects localhost, private-network and non-HTTPS URLs before submission. A model whose input_references descriptor is min 0, max 0 is text-to-image only and rejects references, so read the descriptor from GET /v1/images/models first.
For an edit, send aspect_ratio set to auto to match the reference. Omitting the field is not the same as auto, and the docs say so directly.
The arithmetic
Cost of the run is cost_usd x n per call, and n is 1 here. 31 x $0.15 = $4.65; 31 x $0.1875 = $5.81. Pro at 2K costs $1.16 more across the run, or $0.0375 per edit.
For GPT Image 2.5 the table shows the catalog's medium 2K row, $0.055625, which is a default-options estimate.
| Model | Row per image | 31 edits | Reference inputs |
|---|---|---|---|
| Nano Banana 2.1 2K | $0.15 | $4.65 | 93, not billed separately |
| Nano Banana Pro 2K | $0.1875 | $5.81 | 93, not billed separately |
| GPT Image 2.5 medium 2K (default-options row) | $0.055625 | $1.72 | 93, check usage.cost |
What to verify before you scale
Run one edit per model with your real references and read usage.cost in the response. The docs describe GPT Image 2.5 as token-billed upstream with input image tokens at their own rate, so a call carrying large references can differ from the default-options row. The banana rows are what the catalog publishes per image.
Keep the three references to what the edit needs. Extra references do not change the banana price, but they do make the instruction harder for the model to follow, and that costs you retries.
Choosing between the two banana models
The gap between the models is $1.16 over 31 edits, which is small against the cost of a bad edit that you have to redo. If 2.1 needs 25 percent more attempts per accepted image than Pro, the two cost the same per accepted image, because a retry costs a full extra row.
That is a threshold to test, not a result to assume. Run six edits on each model with the same three references, count the accepted ones, and divide the spend by the accepted count. The docs give you the price side of that ratio; only your own images give you the other side.
Name each image's role in the prompt (garment, model, background) instead of relying on the order of the array, since the docs describe input_references only as an array of image URLs.
Sources
Related posts
More in Developers
- 402 on shot 4 of 6: a $3.00 wallet and a mixed project, step by step
Walk a $3.00 balance through stills, voice, music, six shots and a render. See which submit returns 402, what was already paid, and how to resume safely.
- 402 on the third submit: six 5-second Wan 3.0 jobs on a $1.50 balance
Each 5 s Wan 3.0 720p job reserves $0.625 at submit. On a $1.50 balance the third POST /v1/videos fails with 402 insufficient_credits. Arithmetic and a script.
- 47 voiceover lines, Timeline's 20-part limit: three concats, 3 cents
Timeline audio joins up to 20 parts per job at $0.01. 47 lines need 20 + 20 + 7 = three concats ($0.03). 47 short lines of TTS also bill the 1-cent floor each.
- 48 or 50 fps for YouTube: Sume output.fps accepts 24, 25, 30, 60
YouTube lists 24, 25, 30, 48, 50 and 60 fps as common rates. Sume Timeline's output.fps takes 24, 25, 30 or 60, so omit it for 48 or 50 sources. Why.
Written by Sume