Image to image AI: turn your photo into a new picture
Image-to-image AI redraws your photo from a text instruction: a new style, background or color. What it changes, what it keeps, and how to run it.

Image-to-image AI makes a new picture from a picture you give it plus a text instruction. Instead of starting from words alone, the model takes your image as a reference and redraws it: restyled, with a new background, recolored, or as a variation of the same subject. With Sume's image API, it is the same POST /v1/images call as text-to-image, with your photo's URL in input_references, sent to a model that accepts image input.
Sume facts come from the Image API docs and the model catalog that GET /v1/images/models serves, read on 2026-09-28. The full request, field by field, is in image generation with reference images.
What can image-to-image AI do with my photo?
The prompt says what to change, and the reference supplies what the model should work from. Common jobs:
- Restyle: turn a photo into a painting or an illustration. The docs' own example prompt is "make this scene look like a watercolor painting". More in photo to painting AI.
- Replace the background. The docs' product example reads "Keep the product identical; swap the background to a soft daylight studio". See how to change the background of a photo.
- Recolor one object, such as a jacket, a car or a sofa, as in change the color of an object.
- Make variations: the same subject in a new pose, setting or season.
- Carry one character or product into new scenes by sending the same reference every time, as in consistent character AI image generator.
How is image-to-image different from text-to-image?
Text-to-image starts from the prompt alone. Image-to-image adds one or more images the model must work from, so the result starts from your picture's subject and layout instead of an invented one.
On Sume, a text-to-image-only model refuses a photo with 400 unsupported_parameter, and image generation API models shows how the catalog marks each model. The models below take your photo in input_references, and in current code each one sends it to a dedicated edit mode:
| Model | `model` id | Reference images per request | `aspect_ratio: "auto"` |
|---|---|---|---|
| ChatGPT Image 2.5 | openai/gpt-image-2.5, openai/gpt-image-2.5-sunburst | Up to 16 | Listed |
| ChatGPT Image 2 | openai/gpt-image-2 | Up to 10 | Listed |
| Nano Banana 2 | google/nano-banana-2 | Up to 10 | Listed |
| Nano Banana Pro | google/nano-banana-pro | Up to 10 | Listed |
| Seedream 5.0 Lite | bytedance-seed/seedream-5-lite | Up to 10 | Not listed: send a ratio from its list |
| Seedream 4.5 | bytedance-seed/seedream-4.5 | Up to 10 | Not listed: send a ratio from its list |
How do I run image-to-image with an API?
Put the photo's public HTTPS URL in input_references, say what to change in prompt, and pick a model from the table. The request below is the docs' example with aspect_ratio: "auto" added, which the docs recommend on image-to-image calls so the output matches the reference.
- Localhost, private-network and non-HTTPS URLs are rejected before submission, so host the photo at a public HTTPS address.
- Leaving
aspect_ratioout is not the same as sending"auto". The Seedream models don't listauto, so give them a ratio from their list. - The call waits up to 30 seconds. A
202instead of200means the job isn't finished: pollGET /v1/jobs/{id}/statusuntil it iscompleted,failedorcanceled, and readGET /v1/jobs/{id}/resultonly if it completed.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2",
"prompt": "make this scene look like a watercolor painting",
"input_references": [
{ "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
],
"aspect_ratio": "auto"
}'Will the new image keep my photo's details?
Not exactly. The output is a new image, not your file with one area patched, so faces, small text, logos and fine patterns can drift even when the prompt asks to keep them. Name what must stay the same, change one thing per request, and compare each result with the original. Prompt patterns for this are in AI photo editing prompts.
What does it cost, and what are the limits?
- Each model has a per-image rate in the
pricingline ofGET /v1/images/models/{model_id}/endpoints, plus a 5.5% agent fee by default. On ChatGPT Image 2.5, that price is estimated from tokens, so the output's size and quality and the reference images you send all move it.usage.costin the response is the USD amount billed, and a failed or cancelled generation is not billed. AI image generator API cost lists the rates. - Seedream 4.5 returns one image per call when you send references, whatever
nsays (current code). - Result URLs are Sume-hosted and signed, so download the images you keep.
Sources
Related posts
More in Models
- Image to video prompt examples: what to write
An image-to-video prompt needn't describe the photo again. It says what moves, what the camera does, and what stays still. Examples by photo type.
- Is Kling AI Chinese? Who owns it and where it's based
Yes. Kling AI is developed by Kuaishou Technology, a Beijing-based company listed in Hong Kong. Who owns it, where its API runs, and other access.
- AI video prompt JSON: what the keys actually do
A JSON prompt for AI video is still text: the model reads keys as words. Length, size and audio are request fields, not keys inside the prompt.
- Kling 3.0 API: text or image to video, limits and price
Kling 3.0 has an API: Kling's own, and multi-model APIs such as Sume's, where it is kling-3: 4–15 second clips from text or frames, audio optional.
Written by Sume