Walmart Returns Insights: add a what's-in-the-box image
Walmart's Returns Insights shows return trends. If buyers expected more, make a labeled what's-in-the-box image with Ideogram 4.5 on Sume.

If Walmart's Returns Insights dashboard shows returns from buyers who expected something different, one fix you control is the picture: a what's-in-the-box image with short labels. Ideogram 4.5 on Sume is built for text in images, and you pass your real product photo as the first reference so the item is edited, not reinvented.
Walmart's release notes list Returns Insights on September 11, an enhanced dashboard with return metrics, trends and category recommendations (read 2026-10-04). The notes do not say what causes returns, so read your own return reasons first.
What does Ideogram 4.5 accept on Sume?
From the Image API docs: with no input_references it generates from text; with them it edits the first image and uses up to four more as references, five in total. quality is low, medium or high (medium when omitted), resolution is 1K or 2K, and an edit without aspect_ratio keeps the source image's shape.
| quality | Price per image |
|---|---|
| low | $0.03 |
| medium | $0.06 |
| high | $0.22 |
How do you request the image?
Put the exact label text in the prompt, in quotes, and list only what actually ships in the box. The docs price is the Fal list price, per image, whatever the size.
import os, requests
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json={
"model": "ideogram/ideogram-v4.5",
"prompt": "Flat lay on white: the product, a charging cable, a manual. "
"Clean labels reading 'Speaker', 'USB-C cable', 'Manual'",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/product.jpg"}}
],
"quality": "medium",
"resolution": "2K",
"aspect_ratio": "1:1",
},
timeout=120,
)
print(r.status_code, r.text[:300])
What must you proof?
Read every word in the image. Text models still misspell now and then, and an extra item that is not in the box can create the very mismatch you are trying to prevent. Count the pieces against a real unboxing.
If a label is wrong, re-run with the same references. Our text-edit post covers changing text in an existing picture.
Will it reduce returns?
Nobody can promise that. Compare return reasons on the dashboard before and after the image goes live, for the same SKU over a similar period, and keep the change if the reasons you targeted fall.
Sources
Related posts
More in Use cases
- Walmart rich media is restricted for alcohol, tobacco and firearms
Walmart Marketplace restricts rich media for alcohol, tobacco and firearms and may unpublish it. A category gate to run before you spend on AI clips.
- Avatar script CTA: no click the green button, WCAG 1.3.3
An avatar that says click the green button on the right fails WCAG 1.3.3 for some viewers. How to write the spoken call to action so it still works.
- Autoplay avatar welcome video with sound: WCAG 1.4.2 rule
Can an AI avatar welcome video autoplay with sound? WCAG 1.4.2 allows it only with a pause or volume control once audio runs past 3 seconds.
- Looping avatar video on a page: WCAG 2.2.2 pause rule
A looping avatar clip beside other content must be pausable under WCAG 2.2.2 once it moves for more than 5 seconds. How to meet it with a Sume MP4.
Written by Sume