15 Ideogram 4.5 edits with 5 images each: $1.125 at medium on Sume
Ideogram 4.5 on Sume edits the first image and takes up to four more as references. Fifteen medium edits cost 15 x $0.075 = $1.125.

On Sume, an Ideogram 4.5 edit with five images (the one to edit plus four references) is billed per image output, not per reference. At medium quality that is $0.075, so 15 edits cost 15 x $0.075 = $1.125. At low it is 15 x $0.0375 = $0.5625 and at high 15 x $0.275 = $4.125.
What the docs say about references
The Image API docs say that without input_references Ideogram 4.5 generates from text. With references it edits the first image and uses up to 4 more as references, for 5 in total. Its price is per image by quality for all sizes, and an edit without aspect_ratio keeps the shape of the source image.
| Quality | Per image | Arithmetic | Total |
|---|---|---|---|
| low | $0.0375 | 15 x 0.0375 | $0.5625 |
| medium | $0.075 | 15 x 0.075 | $1.125 |
| high | $0.275 | 15 x 0.275 | $4.125 |
Request shape
Each reference is an input_references entry of type image_url with a public HTTPS URL. Sume rejects localhost, private-network and non-HTTPS URLs before submission. Put the image to edit first, since that is the one Ideogram edits, and the four references after it. Send aspect_ratio: auto if you want the result to match the reference; omitting it is documented as not the same thing.
- Entry 1: the image to edit
- Entries 2 to 5: style or subject references
- aspect_ratio: auto to match the source shape
- Quality low, medium or high sets the price
What to check
Run one edit and read usage.cost to confirm the per-image price. Failed generations are not billed. If you need a mask, Ideogram is not the row: the docs name mask_url for the two GPT Image 2.5 rows. If a file is private, host it behind a public HTTPS URL for the call or choose a source that is.
Budget and ordering
If you are editing product shots, sort the files so the hero image is always first. A reorder changes which image is edited and which are references, and the price does not tell you if you got it wrong. A good habit is to name the reference files with a prefix such as 1-source, 2-style, 3-style and so on, and to sort on the name in code.
For 15 edits at three quality levels, the total range is $0.5625 to $4.125. A single run at medium is $1.125, so a first pass at low and a second at high for the 5 best edits is 15 x $0.0375 + 5 x $0.275 = $0.5625 + $1.375 = $1.9375.
Why the reference count does not change the price
The Sume docs describe Ideogram 4.5 as priced per image by quality for all sizes. That wording is about the output image, and it is the basis for the estimate here. Token-billed rows work differently: for GPT Image 2.5 the docs list input image tokens at $8 per million, so references can add to the cost there. If you move this workflow to a different row, re-measure with one call and read usage.cost.
Sources
Related posts
More in Developers
- 164 days from DALL-E 3 shutdown to GPT Image 1, then 39 more
OpenAI removed DALL-E 3 on May 12, 2026, retires GPT Image 1 on Oct 23, then three more image ids on Dec 1. Day counts and the Sume ids that replace them.
- 26 dialogue lines, a 20-part concat cap: two $0.01 jobs, then offsets
Timeline audio concat takes 1-20 parts. A 26-line dialogue needs two concat jobs, $0.02 total, plus about $0.17 of TTS. How to chain them and keep the offsets.
- A 28-minute talk transcribed on Sume: 3 detach ranges plus STT = $0.31
Audio detach caps output at 900 s and STT estimates cap at 10 minutes. Here is how a 28-minute talk splits into three ranges, and what it costs: $0.31.
- 34 pause cuts, 20 audio parts: split the spine in two levels
Timeline audio.parts holds 20 slices. For 34 pause cuts, build two concat files of 17, feed them as two parts, and re-base each video slot start.
Written by Sume