Launch kit from 6 references: 4 angles, a style card, a logo
ChatGPT Image 2.5 on Sume takes up to 16 references, so one request can carry four product angles, a style card and a logo. How to order and prompt them.

Answer
ChatGPT Image 2.5 on Sume accepts up to 16 reference images, so a single request can carry four product angles, one style card and one logo, six in all, and still leave ten slots free. Name each reference in the prompt by its position, because the list order is the only handle you have.
Reference URLs must be public HTTPS. Sume rejects localhost, private-network and non-HTTPS URLs before the request is submitted.
| Position | Reference | Prompt wording |
|---|---|---|
| 1 | Front view of the product | image 1 is the product, keep its shape and label |
| 2 | Side view | image 2 is the same product from the side |
| 3 | Back view | image 3 is the back, text must stay unchanged |
| 4 | Top view | image 4 is the lid from above |
| 5 | Style card | image 5 sets the palette and lighting only |
| 6 | Logo on white | image 6 is the logo, place it once bottom right |
Request
{
"model": "openai/gpt-image-2.5",
"prompt": "Launch hero: image 1 product centered on a warm stone surface, lighting and palette from image 5, logo from image 6 bottom right",
"quality": "high",
"aspect_ratio": "4:5",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/front.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/style.jpg"}}
]
}Tips
- Use only the references you need. Each one adds input tokens to the estimate, and 2.5 is billed from tokens rather than a flat per-image rate.
- Add
mask_urlwhen you want to change one region and keep the rest. It is optional and specific to the 2.5 rows. background: transparentgives a cut-out for a launch page. If you leave the quality out, it defaults tohigh.- Draft at
lowfirst. Sume's estimator puts a low 1024x768 image at about $0.005 against about $0.045 at high.
The 16-reference ceiling is a catalog value, so confirm it with GET /v1/images/models/openai/gpt-image-2.5/endpoints before you rely on it. Docs: Sume Image API.
Before a large run
Prices and descriptors change when the catalog changes, so confirm them before you spend. Call GET /v1/images/models/{id}/endpoints for the row you plan to use and read its pricing line and supported_parameters; both come back in one response.
Then run a pilot of three to five images and read usage.cost on each response. Multiply by your planned count for a forecast you can trust. Completed generations are billed in full and failed or cancelled ones are not, so a pilot that errors costs nothing.
For big batches, use mode: "async" or mode: "webhook" with a public HTTPS webhook_url, so no request waits on the 30-second sync limit. Poll GET /v1/jobs/{id}/status and fetch GET /v1/jobs/{id}/result when the job completes.
Sources
Related posts
More in Use cases
- Pumpkin patch last-weekend promo clip with a carving-night hook
Use the NRF carving stat to make a 10-second last-weekend promo for a pumpkin patch: two Wan 3.0 shots and a Timeline join, about $0.73 at 480p on Sume.
- Quinceañera invitation teaser: portrait and venue as references
A 10-second quinceañera invitation teaser from up to 10 reference images with Gemini Omni Flash 1.1, with the date added as a plate instead of generated text.
- Re-voice a recording: Sume STT then TTS as two jobs
Transcribe a recording with Sume STT, fix the text, then speak it with Sume TTS. Two jobs, two prices, and where the human edit goes.
- Read a competitor ad before you remix it: 24 frames per call
Sume video-frames pulls up to 24 stills from a clip you own, at times you name, so you can study a reference ad structure before a remix. Limits and errors.
Written by Sume