Pinterest visual search ads: a lifestyle scene from one product photo
Pinterest's visual search ads are a planned beta. Generate a lifestyle image from one product photo with POST /v1/images and an input reference, ratio 2:3.

To prepare images for Pinterest's visual search ads, send your product photo as an input_references item to POST /v1/images and ask for the scene around it, with aspect_ratio: "2:3". A trade tracker lists Pinterest visual search ads as a planned beta (read 2026-10-06), so the format and its rules may still change. I did not verify any Pinterest creative requirement for this format; treat this as a way to make candidate images and read Pinterest's own page before you submit any.
Sume facts are from Image API: generate and edit images; the Pinterest note is from the Social media platform updates, October 2026 (read 2026-10-06).
Why start from the real product photo?
Visual search matches what a shopper sees to what you sell, so the product in your image has to look like the product in the catalogue. Text-only generation invents a plausible mug, not your mug. A reference image conditions the result on your item. The reference must be a public HTTPS URL; Sume rejects localhost, private-network and non-HTTPS URLs before it submits, and a model whose input_references range is 0 to 0 is text-only and rejects references.
How to ask
Describe the scene, not the product. 'On a pale oak kitchen counter, morning window light, a folded linen cloth beside it' gives the model room to draw the setting while the reference keeps the object. Use aspect_ratio: "2:3" for a vertical pin-shaped image, or auto on edits to follow the reference's shape; the API doc notes that omitting the field is not the same as auto. Pick a model from GET /v1/images/models that lists input_references and your ratio, because an unlisted parameter returns 400 unsupported_parameter rather than being dropped.
import json, os, urllib.request
body = {
"model": "openai/gpt-image-2.5",
"prompt": "Place this exact mug on a pale oak kitchen counter, "
"morning window light, a folded linen cloth beside it",
"input_references": [
{"type": "image_url",
"image_url": {"url": "https://media.sume.com/img/demo/mug.png"}}
],
"aspect_ratio": "2:3",
"quality": "high",
}
req = urllib.request.Request(
"https://api.sume.com/v1/images",
data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"},
)
with urllib.request.urlopen(req) as r:
print(r.status, json.load(r))
Reading the response
POST /v1/images waits up to 30 seconds. A 200 is the image response, with data[].url pointing at media.sume.com; a 202 is a job envelope and you poll GET /v1/jobs/{id}/status then fetch /result. Check the status code, not the body shape. Slow settings, such as 4K or high quality, are the most likely to come back as 202.
| Field | Value | Note |
|---|---|---|
| model | openai/gpt-image-2.5 | Supports references and background control |
| input_references | public HTTPS image_url | Up to 16 for this model |
| aspect_ratio | 2:3 or auto | Check the catalog descriptors |
| quality | auto, low, medium, high, xhigh, max | Default high |
| Success | 200 image, 202 job | Branch on status code |
What the model costs
Sume's Image API page lists the token rates behind GPT Image 2.5 (Flare): $30 per million output image tokens, $8 per million input image tokens and $5 per million input text tokens. At 1024 by 1024, xhigh output is $0.09366 and max output is $0.21072, before input tokens and Sume pricing. The reference image counts as input tokens, so an edit costs a little more than the same prompt as text alone. Check the quoted figures against the page before you budget, since pricing sections change.
Because quality defaults to high when you omit it, a first test at medium is a cheaper way to check composition. Raise the quality only for the candidate you will actually use.
What to check before you use the image
Put the original and the result side by side and check the details that identify the product: logo placement, colour, handle shape, label text. Generated scenes sometimes alter small text. If a label changes, retry with the same reference and a calmer prompt, or leave the product photo as is and generate only the backdrop. Also read Pinterest's disclosure rules for AI images; the stored posts on Pinterest's AI label cover that separately.
Sources
Related posts
More in Use cases
- Pinterest visual search ads: a transparent-background product cutout
Visual search ads match your product by its look. Ask for background: transparent on openai/gpt-image-2.5 and keep a clean PNG cutout of each item.
- Podcast network: transcripts, cover art and stings for 12 episodes
12 episodes of 45 minutes, four Ideogram 4.5 cover options each and three music stings each cost about $14.24 a month on Sume. Unit prices and the sum.
- Print-on-demand: 200 designs a month on Ideogram 4.5, plus cutouts
200 text-on-shirt designs a month cost about $16 at Ideogram 4.5 medium or $58 at high on Sume, plus $4.75 for cutouts. The budget, line by line.
- Product video for ecommerce: a white-background photo to 9:16
Turn a white-background product photo into a 5-second vertical scene with POST /v1/videos, seedance-2 and an image_url input reference.
Written by Sume