Zalando PDP video: silent, portrait, no text, AI quality bar
Zalando's PDP video rules: silent, portrait, no text, and no low-quality AI footage. How to check a Sume clip against them before upload.

Zalando's partner guidelines require portrait PDP videos with no sound or written text, and say low-quality AI footage is not permitted. With Sume you can request a portrait clip from POST /v1/videos, confirm there is no audio track with Video inspect, and pull a thumbnail with Video frames.
The rules as summarized
Rules are as read on the Zalando Partner University page on 2026-10-03; they can change, so confirm each line before you submit.
| Item | Reported rule |
|---|---|
| Start | Uploadable since Nov 2025 |
| Delivery | Merchant Product Submission API or Article XML feeds |
| Orientation | Portrait, ratio 1.44 to 1.8 |
| Content | Written text and sound are not permitted |
| Format | MP4, H.264, at least 762x1100 px, 24+ fps, 2000+ Kb/s, up to 250 MB |
| Thumbnail | One JPEG preview image, at least 762x1100 px |
| Quality | Pixelation, motion blur, unnatural textures and anatomical inconsistencies are not permitted |
Make it portrait
Sume's aspect ratios include 9:16, 2:3, 3:4 and others, and each model lists the subset it accepts in supported_aspect_ratios. If the Zalando ratio is long side divided by short side, 9:16 is about 1.78 and 2:3 is 1.5, both inside the range, while 3:4 is about 1.33, which is outside it. Confirm how Zalando defines the ratio, and check that your chosen model offers the ratio you need.
Make it silent
POST /v1/videos has a generate_audio field that defaults to the model's audio capability. Not every model can switch audio off: Gemini Omni Flash 1.1 always has native audio and rejects generate_audio: false. Read generate_audio in GET /v1/videos/models and pick a model that allows it.
Then verify the output rather than trusting the setting. Video inspect with frames: false returns probe facts, including probe.has_audio.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: pdp-check-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/pdp.mp4", "frames": false}'The thumbnail and the quality bar
Video frames extracts stills at times you name from one hosted clip, as durable image artifacts at source size. Pass at[] with a timestamp to pick a clean frame as the thumbnail. It is billed by its Modal compute.
The quality rule is the hard part. No setting guarantees that a clip is free of motion blur or anatomy errors. Look at the stills from Video inspect, reject clips with the defects Zalando lists, and regenerate with a shorter, simpler prompt. Keep a human check in the loop before upload.
Sources
Related posts
More in Use cases
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
- AI avatar for YouTube videos: Shorts and long-form
Use an AI avatar in YouTube videos: a 9:16 talking video of up to 60 seconds for a Short, or 16:9 segments joined into one long-form video.
Written by Sume