GPT Image 2.5 character drift: what OpenAI says and what to send
OpenAI's image guide says GPT Image may struggle to keep a recurring character the same across generations. What to send on Sume to reduce drift.

Yes, GPT Image can drift. OpenAI's image generation guide lists consistency as a limitation: the model can produce consistent imagery but "may occasionally struggle" to keep recurring characters or brand elements the same across multiple generations. The practical fix is to stop describing the character from scratch each time and send the same reference image with every request, which on Sume means input_references on POST /v1/images.
OpenAI's statement is from its Image generation guide, read 2026-10-02. Sume's request fields are from the Sume Image API page, read the same day.
What exactly does OpenAI say?
The guide's Limitations list has four items: latency (complex prompts may take up to 2 minutes), text rendering, consistency, and composition control. The consistency item says that while the model can produce consistent imagery, it may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations.
That is a vendor caveat, not a measurement. It does not say how often drift happens, so treat any single series as something to check by eye.
Which Sume request reduces drift?
Use the same reference image in every call and put the unchanging identity words in one reusable block. Sume's catalog advertises up to 16 input_references for openai/gpt-image-2.5 and openai/gpt-image-2.5-sunburst, so a face shot, an outfit shot and a product shot can all ride along.
Edits behave differently from fresh generations: the docs say to prefer aspect_ratio: "auto" on edit calls so the output follows the reference, and that leaving it out is not the same thing.
| Hold fixed | How on Sume | Why |
|---|---|---|
| Face and outfit | Same image in input_references | The model sees the same pixels each time |
| Identity wording | One prompt block, pasted unchanged | Only the scene changes between calls |
| Frame shape | aspect_ratio: "auto" on edits | Output follows the reference shape |
| Model id | One pinned id such as openai/gpt-image-2.5 | Avoids a different model mid-series |
Does Sume keep state between calls?
No. Each POST /v1/images call carries its own prompt and references, and the Image API page documents no conversation or response-id field. OpenAI offers multi-turn editing in its Responses API, which Sume's Image API does not mirror. Re-send the reference on every call.
Also note that seed is in the Sume schema but no model advertises it, so a request that sends one returns 400 unsupported_parameter. You cannot pin a seed to get repeatable characters.
How do I check a series for drift?
Generate the frames, put them side by side, and compare the parts that must not move: face shape, hair, logo placement. If frame four drifts, re-run it with the first frame as the reference rather than the third, so errors do not stack up from one edit to the next.
A slow high-quality frame may return 202 with a job envelope; read the images from GET /v1/jobs/{id}/result as the Jobs and results page describes.
Sources
Related posts
More in Models
- gpt-image-2.5 smallest custom size: the 655,360-pixel floor
gpt-image-2.5 on Sume rejects images under 655,360 pixels. Why 512x512 and 768x768 fail, which small sizes pass, and the other edge rules to check first.
- Grok Imagine video in Sume: why it needs a start-frame image
Sume's Grok Imagine entry is image-to-video only: it blocks a submit without a start frame, tops out at 10 seconds, and sends no audio or aspect ratio.
- H3 Max 3D to Video: previs to photoreal on fal, not on Sume
fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.
- H3 Max Insert-Video: add a scene mid-clip on fal, and on Sume
fal's H3 Max Insert-Video adds 5 to 13 s to a source up to 60 s, billed on the new seconds only. Sume lists no such row; here is a trim, generate, join route.
Written by Sume