Zalando required images per article: 3, 2 or 1, then a batch plan
Zalando needs 3 images for apparel, accessories and shoes, 2 for home, 1 for beauty and others. A per-SKU plan for generating the missing views on Sume.

Zalando's required image count depends on the category: 3 for apparel (including kids), accessories and shoes, 2 for home, and 1 for hosiery, equipment, beauty, toys and electronics, with a catalogue view in every case. Jewellery, per its own guide, needs a primary packshot plus two further compliant images. These come from the image guidelines, updated October 1, 2026, and the category guides (read 2026-10-02). For a catalogue, that turns into a list of views per SKU, and one Sume job per view.
How many images does each category need?
Duplicates do not count, and a multipack must include a primary packshot that shows all items in the pack. Electronics and toys need one image showing the CE label on the device or packaging, so that one is a real photo of a real label, not something to generate.
| Category | Total required | Catalogue view |
|---|---|---|
| Apparel, accessories, shoes | 3 | Model front crop or primary packshot (shoes and jewellery: primary packshot only, per their guides) |
| Home | 2 | Primary packshot |
| Hosiery | 1 | Model front crop or primary packshot |
| Equipment, beauty, toys, electronics | 1 | Primary packshot |
| Jewellery (own guide) | 3 | Primary packshot, plus two further compliant images |
Which views can I derive from one photo?
That depends on what the photo shows. A reference-guided edit can change a background or remove a mannequin, but a back view or a detail of a fastener needs to be visible in the source or photographed. Zalando's jewellery guide asks for a back view where the fasteners can be seen, and its shoes guide asks for sole and detail views as additional views. A model cannot invent a clasp it has never seen without risking a product that looks different from the real one.
So split the plan in two: views you can edit from existing photos, and views you must shoot.
How do I batch it on Sume?
Build a list of (SKU, view) pairs, send one POST /v1/images per pair with input_references pointing at the right source photo, image_size of 2000x2880, and output_format of jpeg, and give every request its own Idempotency-Key such as sku-1042-front. The docs recommend an idempotency key on every paid submit that may be retried (Generation admission).
Each call either returns 200 with the image or 202 with a job to poll. Store the job id with the SKU so a reviewer can trace each file. Prices differ per model, so read them from GET /v1/images/models and multiply by your view count before you start.
What should I do next?
Here is the short version, in the order to do it.
- List the category of each SKU and its required count.
- Mark which views exist in photos already.
- Generate only the derivable ones, one job per view.
- Review each output against the real product before upload.
Sources
- Zalando Partner: Zalando image guidelines, updated October 1, 2026 (read 2026-10-02)
- Zalando Partner: Apparel image guide, updated March 3, 2025 (read 2026-10-02)
- Zalando Partner: Shoes image guide, updated July 4, 2025 (read 2026-10-02)
- Zalando Partner: Jewellery image guide, updated March 3, 2025 (read 2026-10-02)
- Sume docs: Image API
- Sume docs: Generation admission
Related posts
More in Use cases
- Zalando shoe packshot: left shoe pointing left, laces tied, centred
Zalando's shoes guide: primary packshot as catalogue view, left shoe pointing left, laces tied, centred. How to write the edit prompt on Sume.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
Written by Sume