GPT Image 2.5 mask edit, then an Ideogram 4.5 text pass: one chain

Use GPT Image 2.5 for the masked region change and Ideogram 4.5 for the text pass, both on POST /v1/images. When it beats one model, and what each step bills.

5 min readSume
All posts

Use GPT Image 2.5 first when you need to change a region with a mask, and Ideogram 4.5 second when you need to fix the text; both are called through POST /v1/images, and the output URL of step one is the first input_references entry of step two. It is worth the extra call only when one model alone cannot do both jobs.

The split follows what Sume's catalog allows. mask_url is accepted only on GPT Image 2.5; every other image model, Ideogram 4.5 included, answers 400 unsupported_parameter. Ideogram calls 4.5 "the most precise edit model" (Ideogram on X, read 2026-10-05), which is a claim about editing precision, not about masks.

What does each model contribute?

The table lists the differences that matter for a two-step chain. Everything in it is from the Sume Image API docs and catalog, read 2026-10-05.

GPT Image 2.5 and Ideogram 4.5 on Sume, read 2026-10-05.
FieldGPT Image 2.5Ideogram 4.5
Public idopenai/gpt-image-2.5ideogram/ideogram-v4.5
mask_urlSupported400 unsupported_parameter
backgroundSupportedNot listed
References per requestup to 165 (first is edited)
Price basisToken-priced, see the docs$0.0375 low, $0.075 medium, $0.275 high

What does the chain look like?

Step one sends the source, the mask and a prompt about the region only. Step two sends step one's result with a prompt that quotes the text to change. If both steps use PNG and the same shape, nothing needs resizing between them. Without aspect_ratio, an Ideogram edit keeps the source geometry, so step two will not reshape what step one made.

  • Step 1, GPT Image 2.5: input_references: [source], mask_url, prompt "Replace the sofa with a green velvet sofa".
  • Check: run a pixel diff outside the mask; see verify the edit stayed inside the region.
  • Step 2, Ideogram 4.5: input_references: [step1_url], prompt "Change the sign text from OPEN to SALE. Keep everything else unchanged.".
  • Check: read the sign at 100 percent and compare the rest to the step-one output.

When is one model enough?

If the change is text only, skip step one. Ideogram 4.5 alone takes it for $0.0375 at low. If the change is a region swap with no text to fix, GPT Image 2.5 alone does it with the mask and a second model adds only cost. The chain pays off when a masked change leaves new text behind, like a swapped product with a wrong label, and you want the region boundary enforced by the mask and the lettering by the text-focused model.

What no page in the docs promises is that the second model will keep step one's pixels unchanged. Treat that as untested until you diff your own pair of images.

What does the chain bill?

Step two is a fixed price by quality, and step one depends on the GPT Image 2.5 pricing, which is per token and varies with output size and quality. The docs give list prices for the 1024 output only, so read usage.cost from step one's response and add the Ideogram price for your tier. The cost of 1,000 images across the three models has the per-image comparison.

Keep one Idempotency-Key per step, and do not start step two until step one returns 200, or finishes as completed if it came back 202. Save step one's output as PNG before using it as an input.

How do I decide before I build it?

Run the chain by hand on five of your own images and write down, for each, what the first step got right and what the second step changed. If the second step never fixes anything the first missed, drop it. If it fixes text every time, keep it and automate it.

Also check that step two did not undo step one. Diff the region step one changed against the same region after step two; if it moved, you have a two-model chain that fights itself, and the answer is a single model with a stricter prompt.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume