GPT Image 2.5 character drift: what OpenAI says and what to send

OpenAI's image guide says GPT Image may struggle to keep a recurring character the same across generations. What to send on Sume to reduce drift.

5 min readSume
All posts

Yes, GPT Image can drift. OpenAI's image generation guide lists consistency as a limitation: the model can produce consistent imagery but "may occasionally struggle" to keep recurring characters or brand elements the same across multiple generations. The practical fix is to stop describing the character from scratch each time and send the same reference image with every request, which on Sume means input_references on POST /v1/images.

OpenAI's statement is from its Image generation guide, read 2026-10-02. Sume's request fields are from the Sume Image API page, read the same day.

What exactly does OpenAI say?

The guide's Limitations list has four items: latency (complex prompts may take up to 2 minutes), text rendering, consistency, and composition control. The consistency item says that while the model can produce consistent imagery, it may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations.

That is a vendor caveat, not a measurement. It does not say how often drift happens, so treat any single series as something to check by eye.

Which Sume request reduces drift?

Use the same reference image in every call and put the unchanging identity words in one reusable block. Sume's catalog advertises up to 16 input_references for openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst, so a face shot, an outfit shot and a product shot can all ride along.

Edits behave differently from fresh generations: the docs say to prefer aspect_ratio: "auto" on edit calls so the output follows the reference, and that leaving it out is not the same thing.

What to hold fixed between frames of a series; the Sume fields are from the Image API page, read 2026-10-02.
Hold fixedHow on SumeWhy
Face and outfitSame image in input_referencesThe model sees the same pixels each time
Identity wordingOne prompt block, pasted unchangedOnly the scene changes between calls
Frame shapeaspect_ratio: "auto" on editsOutput follows the reference shape
Model idOne pinned id such as openai/gpt-image-2.5Avoids a different model mid-series

Does Sume keep state between calls?

No. Each POST /v1/images call carries its own prompt and references, and the Image API page documents no conversation or response-id field. OpenAI offers multi-turn editing in its Responses API, which Sume's Image API does not mirror. Re-send the reference on every call.

Also note that seed is in the Sume schema but no model advertises it, so a request that sends one returns 400 unsupported_parameter. You cannot pin a seed to get repeatable characters.

How do I check a series for drift?

Generate the frames, put them side by side, and compare the parts that must not move: face shape, hair, logo placement. If frame four drifts, re-run it with the first frame as the reference rather than the third, so errors do not stack up from one edit to the next.

A slow high-quality frame may return 202 with a job envelope; read the images from GET /v1/jobs/{id}/result as the Jobs and results page describes.

Sources

Related posts

More in Models

All Models posts

Written by Sume