Product photo into a new scene: Ideogram 4.5, one packshot, four moods
With references, Ideogram 4.5 on Sume edits the first image and takes up to four more as references. Use that order to put one packshot into a new scene.

To put one product photo into a new scene with ideogram/ideogram-v4.5, send the packshot as the first entry of input_references and up to four mood images after it. The Sume image docs say that with references the model edits the first image and uses up to four more as references, five in total. The order is the whole technique: the first image is what gets edited, and the rest only steer it. Ideogram describes 4.5 as its most precise edit model, built so that multi-turn editing works (read 2026-10-05).
Order the references
Write the list in the order you want the model to treat the images.
- Position 1: the packshot, shot square-on with the product complete in frame.
- Position 2: a surface or room you want, with no product in it.
- Position 3: a light reference, such as a window or a hard flash look.
- Position 4: a color or props reference.
The request
quality takes low, medium or high and defaults to medium, and resolution takes 1K or 2K. An edit without aspect_ratio keeps the shape of the source image, so leave it out when the output must match the packshot, and set aspect_ratio only when you want a new crop. Every reference URL must be public HTTPS, and Sume rejects localhost, private-network and non-HTTPS URLs before it submits.
The route uses the usual mode field, and sync is the default on this route. Add metadata with your SKU: Sume stores it on the job and does not send it to the provider.
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sku-4471-scene-v1" \
-d '{
"model": "ideogram/ideogram-v4.5",
"prompt": "Place the product from the first image on the oak table from the second image, in the light of the third. Keep the label unchanged.",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/sku-4471-packshot.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/oak-table.jpg"}},
{"type": "image_url", "image_url": {"url": "https://example.com/window-light.jpg"}}
],
"quality": "high",
"resolution": "2K"
}'Price and what you cannot know in advance
The docs give the Fal list price for Ideogram 4.5 as $0.03, $0.06 or $0.22 per image by quality, for all sizes, before the Sume margin. Run at low while you tune the references, then move to high for the keepers. A misspelled label is still a completed, billed job, so check the text on the product in each result before you ship it.
If a scene is nearly right, do not restart from the packshot. Take the result and send it as the first reference of the next edit with one change in the prompt. Ideogram's own post says the model is built to avoid artifact buildup across turns, which is the property that makes this safe for two or three rounds.
Run one SKU through three scenes
A useful test is one packshot against three scene sets: kitchen, desk and outdoor table. Keep the packshot in position 1 for every request and change only the other references and one clause of the prompt. Give each request its own Idempotency-Key, such as the SKU plus the scene name, so a retry of the kitchen job cannot be mistaken for the desk job. Compare the three results for the same three failure points: label text, product edge against the new surface, and the shadow direction against the light reference.
If the shadow is wrong, add a sentence to the prompt that names the light direction, and do not add more references. If the label text drifts, lower the number of references to two, because fewer steering images leave less room to rewrite the product.
When to use another model
For a hard mask, where only the table should change and every pixel of the product must stay, use ChatGPT Image 2.5 with a mask_url. It takes up to 16 image references and a background of auto, transparent or opaque. Ideogram has no mask parameter on the Sume page, so a mask job belongs to the other model.
Sources
Related posts
More in Use cases
- Product photo to video: first frame or reference for each SKU?
Use frame_images when the photo must open the clip, input_references when it should guide the look. Send both and Sume runs image-to-video.
- Recall notice as a 9:16 and 16:9 avatar video: cost and captions
A recall notice needs to reach people on phones and on a site. Two 30-second Sume avatar jobs, one per aspect ratio, cost $14.70 at plus quality with captions.
- Proof watermark for client drafts of AI images, in Pillow
Stamp a tiled diagonal PROOF text over an image from Sume before sending drafts. Layer, rotate, alpha-composite, and why it is a courtesy and not security.
- Proposal walkthrough video with an AI avatar: 40 seconds, by tier
Turn a sales quote into a 40-second avatar walkthrough: scope, timeline and price in three scenes, the request body, and the per-quote cost on each Sume tier.
Written by Sume