How to make an unboxing video with AI, step by step
An unboxing video shows hands opening the package and revealing the product. How to make one with AI from a box photo and a product photo.

To make an unboxing video, show hands opening a product's package and revealing what is inside, in a few short beats: the sealed box, the opening, the product coming out, and the product in hand or in use. With AI, you can generate those beats from a photo of the box and a photo of the product instead of filming them. The result is a generated scene, not footage of your actual item.
The Sume steps below come from the Format catalog, Create a run and Video generation docs, read on 2026-09-28.
What is an unboxing video?
It is a short clip of a product's package being opened and its contents revealed. The point is the reveal: the viewer sees what arrives and how it is packed before seeing the product used. Some unboxings have a presenter on camera; a version that shows only the hands and the package also works as faceless UGC.
Which way should I make it with AI?
There are three ways on Sume, from least to most control. Start with the first; move down only if you need to set each shot yourself.
| Way | You send | You decide | More detail |
|---|---|---|---|
Catalog Format sume-product-usage-demo | Box and product photos in attachments, the beats in instruction | The brief; the Format's tools pick the models | Below |
One reveal clip on POST /v1/videos | The closed box as first_frame, the product out as last_frame | Model, prompt, duration, aspect ratio, resolution | Product teaser video |
| Several clips joined into one video | One clip per beat, plus a voiceover; in current code a Timeline 1.0 join keeps only the voiceover and an optional soundtrack, not each clip's own sound | Each beat's order and length | AI product demo video API |
How do I make one with the product demo Format?
The catalog Format sume-product-usage-demo is described as making a product-usage video that "demonstrates one real action, texture, or result in a casual social setting", and it lists "unbox-and-use clips" among its uses. Send a photo of the closed box and one of the product as attachments (public HTTPS URLs, up to 30 images), and write the beats you want in instruction. Read the Format's description first with GET /v1/formats/sume/sume-product-usage-demo.
curl -sS -X POST "https://api.sume.com/v1/formats/sume/sume-product-usage-demo/runs" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mug-unboxing-v1" \
-d '{
"instruction": "Vertical unboxing clip. Hands open the attached box, lift out the mug, and pour a coffee into it. No face on camera.",
"attachments": [
{ "type": "input_image", "image_url": "https://example.com/mug-box.jpg" },
{ "type": "input_image", "image_url": "https://example.com/mug.jpg" }
],
"generation_spend_cap_usd": 20
}'Can I set the first and last shot myself?
Yes, with one image-to-video clip: the sealed box as the first_frame and the product out of the box as the last_frame in frame_images, so the model generates the opening between them. Only models whose supported_frame_images lists last_frame take an end frame, and in current code a last_frame without a first_frame is refused. The product teaser post walks through the same reveal call, and first and last frame lists which models take which frame.
How much does an AI unboxing video cost?
A Format run's generation is metered at the API pricing rates, plus a 5.5% agent fee by default, and generation_spend_cap_usd caps it per run, up to $500. A single clip on /v1/videos is reserved on submit at the provider's list price × 1.25, and usage.cost on the job is the billed amount. Each model's rates are in pricing_skus from GET /v1/videos/models.
What should I check before I post it?
- The packaging: check the print, logos and colors on the box in every beat. Product logo warping shows how to check stills.
- The contents: count what comes out of the box against what really ships in it.
- Where a marketplace or platform asks for footage of the actual item, film it; a generated unboxing is not that.
- Sound: a clip only has an audio track if the model generates one (
generate_audio); otherwise add a voiceover.
Sources
Related posts
More in Use cases
- How to create meeting minutes from an audio recording
Create meeting minutes from a recording in two steps: transcribe the audio, then have a language model draft the summary, decisions, and action items.
- Create motivational videos with AI: voice, shots, and music
Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.
- Press release video: an announcement read by an AI avatar
A press release video reads the headline, key facts, and a quote in about a minute, captioned from the release text. How to make one with an AI avatar.
- New product launch video with an AI avatar presenter
A new product launch video names the problem, shows what's new, and ends on one call to action. Make it from launch copy with an AI avatar presenter.
Written by Sume