Subscription box reveal from six item photos: Wan references, $1.25
Wan 3.0 takes up to 10 reference images: six item photos make a 10 s box reveal at 720p for $1.25 on Sume. References guide the look; they are not exact frames.

A subscription-box or gift-set reveal can be made from up to ten item photos in one Wan 3.0 request. Send each item as an input_references image, ask for a box opening with the items appearing in turn, and a 10-second 720p clip costs $1.25 on Sume (read 2026-10-05: $0.125 a second billed). The catch matters for a box: the Sume videos docs say reference images give "visual guidance, not accurate frames", so labels and logos may not survive.
The request
Reference-to-video uses input_references; image-to-video uses frame_images. Per the docs, if you send both, frame_images controls the mode, so pick one mode per request. The Wan row in the Sume catalog allows up to 10 reference images, 2 to 30 s, and 9:16 among its aspect ratios.
The Python below builds the body and checks the reference limit before you pay. It only prints the body; send it to POST /v1/videos with your key and an Idempotency-Key.
import json
ITEMS = ["https://media.sume.com/artifacts/demo/item%d.png" % n for n in range(1, 7)]
body = {
"model": "wan-3.0",
"prompt": "Top-down reveal of a gift box opening, each item appearing in turn",
"input_references": [
{"type": "image_url", "image_url": {"url": url}} for url in ITEMS
],
"duration": 10,
"resolution": "720p",
"aspect_ratio": "9:16",
}
assert len(body["input_references"]) <= 10, "Wan 3.0 takes at most 10 reference images"
print(json.dumps(body)[:120])
print("billable USD:", 10 * 0.125)
Pricing the box
The clip is the main cost. Add captions only if you burn in text: the caption job is $0.20 for a video of 60 seconds or less (caption docs).
| Length | Clip price | With $0.20 captions | Notes |
|---|---|---|---|
| 5 s | $0.625 | $0.825 | Two or three items |
| 10 s | $1.25 | $1.45 | Six items, one reveal each |
| 15 s | $1.875 | $2.075 | Up to ten items |
What to do about label drift
Because references are guidance, treat the clip as a mood piece and handle exact text separately. Keep product names and the box brand out of the prompt and burn them in with caption cues afterwards, which render your exact string. Check each item against its photo before you publish.
If an item must appear exactly as shot, use the photo as a first frame instead and animate that single item; a full box can then be assembled from several clips in Timeline. The gift-set hero clip guide takes that single-clip route with a Format.
- List the photos in the order you want items revealed and say so in the prompt.
- Use 10 s for six items, which leaves about a second and a half per item.
- Regenerate once at most; a second retake costs as much as the first clip.
For the limits of references across the Wan, H3 and Omni rows, see the reference-clip limits side by side.
Sources
Related posts
More in Use cases
- Audio transcript to timestamped Markdown notes: Python with Sume STT
Turn a 10-minute recording into Markdown with [mm:ss] sentence lines for Notion or Obsidian using Sume STT sentence segments. Cost is about 10 cents per file.
- Suno v6 on licensed data: does it change brand commercial-use risk
Suno v6 is reported to use licensed training data, but suits continue and commercial rights still need a paid plan. What changes for a brand and what does not.
- Suno retires older models: keeping a brand jingle reproducible
TechCrunch says Suno will retire older models after v6. A brand jingle needs a stored master file and a saved prompt, not a promise to regenerate it.
- Swap the person in a clip: reference video or H3 Max Recast on Sume
Seedance 2.5 offers reference-based editing. For a straight person swap, Sume lists H3 Max Recast. When to use which, and the inputs each needs.
Written by Sume