Zalando shoe packshot: left shoe pointing left, laces tied, centred
Zalando's shoes guide: primary packshot as catalogue view, left shoe pointing left, laces tied, centred. How to write the edit prompt on Sume.

For Zalando shoes, the catalogue view has to be a primary packshot; a front crop is not allowed for that category. The guide says to centre the article vertically and horizontally, always show the left shoe pointing to the left, tie laces as if the shoe is worn, avoid overexposure and remove dust. The rules are in Zalando's Shoes image guide, updated July 4, 2025 (read 2026-10-02). Writing those into an edit prompt is the part you control on Sume.
What does the shoes guide require?
Shoes need one compliant primary packshot as the catalogue view and two further compliant images. Duplicating images to reach the count leads to rejection. Side, back and top views are compliant when they show only one shoe. Sole, detail and back views are allowed only as additional views, because the article must be fully recognisable, and they do not count towards the required number.
Run the same prompt across a batch of SKUs only after you have passed three or four by eye. The same wording can behave differently on a boot, a sandal and a trainer, and Zalando adds category rules: for mid-calf or higher shoes its model views crop below the knee, and above-knee boots crop mid-thigh. Packshots are different from model views, so keep the two prompts separate and check each against the guide.
| Rule | What the page says |
|---|---|
| Catalogue view | Primary packshot; front crop not allowed for shoes |
| Primary packshot position | Centred vertically and horizontally |
| Orientation | Left shoe pointing to the left |
| Laces | Tied as if the shoe is being worn |
| Quality | Avoid overexposure; remove dust |
| Sole, detail, back views | Additional views only; do not count to the minimum |
How do I prompt the edit?
Start from a real photo of the shoe and say what must not change: shape, colour, stitching, sole, logo, lace pattern. Then list the packshot rules as instructions: plain white background, shoe centred, left shoe pointing left, laces tied, no text. Sume's docs note that on edit calls you should prefer aspect_ratio: "auto" to match the reference, but when you set image_size it takes precedence, so use image_size for the Zalando 1:1.44 slot (see the size post).
PROMPT = "Edit this photo into a product packshot of the same shoe. Keep shape, colour, stitching, sole, logo and laces exactly as in the photo. Show the left shoe pointing to the left, laces tied, centred, plain white background, soft even light, no text, no extra objects."
import os, requests
r = requests.post(
"https://api.sume.com/v1/images",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "zalando-sku-1042-primary",
},
json={
"model": "openai/gpt-image-2.5",
"prompt": PROMPT,
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/sku-1042.jpg"}}
],
"image_size": "2000x2880",
"output_format": "jpeg",
},
timeout=60,
)
if r.status_code == 200:
print(r.json()["data"][0]["url"])
elif r.status_code == 202:
print("poll", r.json()["data"]["status_url"])
else:
r.raise_for_status()What can go wrong?
Pointing direction is the likely failure: a model may mirror the shoe or draw the wrong side. Look at the result, and if the shoe is wrong, re-run with the source photo that already faces the right way rather than asking for a flip, because a flip can mirror a logo. Zalando also says it may reject AI-generated assets that fail to represent the product accurately, so check heel shape, sole thickness and branding against the real pair.
What should I do next?
Here is the short version, in the order to do it.
- Pick source photos where the left shoe already points left.
- Generate the primary packshot first, then the two further views.
- Check logos and soles against the real pair.
- Do not count sole or detail views toward the minimum.
Sources
Related posts
More in Use cases
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
- AI avatar for YouTube videos: Shorts and long-form
Use an AI avatar in YouTube videos: a 9:16 talking video of up to 60 seconds for a Short, or 16:9 segments joined into one long-form video.
Written by Sume