GPT Image 2.5 mask edit, then an Ideogram 4.5 text pass: one chain
Use GPT Image 2.5 for the masked region change and Ideogram 4.5 for the text pass, both on POST /v1/images. When it beats one model, and what each step bills.

Use GPT Image 2.5 first when you need to change a region with a mask, and Ideogram 4.5 second when you need to fix the text; both are called through POST /v1/images, and the output URL of step one is the first input_references entry of step two. It is worth the extra call only when one model alone cannot do both jobs.
The split follows what Sume's catalog allows. mask_url is accepted only on GPT Image 2.5; every other image model, Ideogram 4.5 included, answers 400 unsupported_parameter. Ideogram calls 4.5 "the most precise edit model" (Ideogram on X, read 2026-10-05), which is a claim about editing precision, not about masks.
What does each model contribute?
The table lists the differences that matter for a two-step chain. Everything in it is from the Sume Image API docs and catalog, read 2026-10-05.
| Field | GPT Image 2.5 | Ideogram 4.5 |
|---|---|---|
| Public id | openai/gpt-image-2.5 | ideogram/ideogram-v4.5 |
mask_url | Supported | 400 unsupported_parameter |
background | Supported | Not listed |
| References per request | up to 16 | 5 (first is edited) |
| Price basis | Token-priced, see the docs | $0.0375 low, $0.075 medium, $0.275 high |
What does the chain look like?
Step one sends the source, the mask and a prompt about the region only. Step two sends step one's result with a prompt that quotes the text to change. If both steps use PNG and the same shape, nothing needs resizing between them. Without aspect_ratio, an Ideogram edit keeps the source geometry, so step two will not reshape what step one made.
- Step 1, GPT Image 2.5:
input_references: [source],mask_url, prompt "Replace the sofa with a green velvet sofa". - Check: run a pixel diff outside the mask; see verify the edit stayed inside the region.
- Step 2, Ideogram 4.5:
input_references: [step1_url], prompt "Change the sign text from OPEN to SALE. Keep everything else unchanged.". - Check: read the sign at 100 percent and compare the rest to the step-one output.
When is one model enough?
If the change is text only, skip step one. Ideogram 4.5 alone takes it for $0.0375 at low. If the change is a region swap with no text to fix, GPT Image 2.5 alone does it with the mask and a second model adds only cost. The chain pays off when a masked change leaves new text behind, like a swapped product with a wrong label, and you want the region boundary enforced by the mask and the lettering by the text-focused model.
What no page in the docs promises is that the second model will keep step one's pixels unchanged. Treat that as untested until you diff your own pair of images.
What does the chain bill?
Step two is a fixed price by quality, and step one depends on the GPT Image 2.5 pricing, which is per token and varies with output size and quality. The docs give list prices for the 1024 output only, so read usage.cost from step one's response and add the Ideogram price for your tier. The cost of 1,000 images across the three models has the per-image comparison.
Keep one Idempotency-Key per step, and do not start step two until step one returns 200, or finishes as completed if it came back 202. Save step one's output as PNG before using it as an input.
How do I decide before I build it?
Run the chain by hand on five of your own images and write down, for each, what the first step got right and what the second step changed. If the second step never fixes anything the first missed, drop it. If it fixes text every time, keep it and automate it.
Also check that step two did not undo step one. Diff the region step one changed against the same region after step two; if it moved, you have a two-model chain that fights itself, and the answer is a single model with a stricter prompt.
Sources
Related posts
More in Comparisons
- GPT Image 2.5 medium is 6x cheaper than Nano Banana 2 1K on Sume
On Sume, GPT Image 2.5 at medium costs $0.0165 for 1024x1024 and Nano Banana 2 1K costs $0.10, a 6.1x gap. Ratios at 2K and 4K, and what the price omits.
- Griffin clones a voice from about 10 seconds: what Sume does for voice
Tavus says Griffin can clone a voice from about 10 seconds of audio. Sume's TTS docs describe choosing a voice, not training one. Plan around that gap.
- Griffin-Lite or a rendered avatar clip for Q4? A decision table
Tavus Griffin-Lite is a research preview for invited testers. Sume Avatar 1.0 renders scripted clips from $11.04 per minute. Pick by the job, not the demo.
- Griffin-Lite vs Sume Avatar 1.0: what you can integrate this week
Compare what Tavus Griffin-Lite shows on its page with what Sume Avatar 1.0 ships in its API: access, input, output, latency claims, tiers and disclosure.
Written by Sume