Gift Unwrap Reveal Video: Wrapped-Box Photo to Product Photo

Turn a photo of a wrapped box and a photo of the product into an unwrap reveal: first frame, last frame, which Sume models take both, and what to check.

4 min readSume
All posts

An unwrap reveal is the cleanest use of a last frame. Send the wrapped-box photo as the first frame and the product photo as the last frame, describe the unwrapping in the prompt, and the model has to travel between two pictures you control. The end of the clip is then your real product, not the model's idea of one.

Read on 2026-10-03: the Sume Videos docs, Sume's catalog constraints in the repository, and Google's Veo page.

Which models take a first and a last frame?

On POST /v1/videos, frame_images entries carry frame_type of first_frame or last_frame. A last frame without a first frame is rejected with unsupported_capability. Grok Imagine Video 1.5 has no end frame on Sume, so it cannot do this.

  • Google says the same for Veo: lastFrame must be used with the image parameter, and 1080p and 4K need 8 seconds.
  • At a list rate of $0.10 per second for Omni at 720p (2026-08-28), a 6-second clip is $0.60 before Sume's margin.
End-frame support on Sume, catalog constraints read 2026-10-03
ModelFirst and last frameDurationAudio
kling-3Yes4 to 15 sOptional; priced separately on and off
wan-3.0Yes2 to 30 sSupported
minimax-h3Yes5 to 15 sAlways on
gemini-omni-flash-1.1Yes3 to 10 sAlways on
grok-imagine-video-1.5No4 to 15 sNo toggle

Shoot the two frames to match

Photograph the wrapped box and the product from the same camera position, on the same surface and in the same light. The box should be roughly where the product will end up, so the model does not have to move the camera and rebuild the room at once. Then write the prompt as an action: the ribbon loosens, the paper folds away off-frame, the lid lifts, the product is left centered.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def frame(url, kind):
    return {"type": "image_url", "image_url": {"url": url}, "frame_type": kind}
body = {
    "model": "gemini-omni-flash-1.1", "duration": 6, "resolution": "720p",
    "aspect_ratio": "9:16",
    "prompt": "Hands pull the ribbon loose, the paper folds away, the lid lifts and "
              "the candle is left centered. Fixed camera, soft light.",
    "frame_images": [frame("https://example.com/box-wrapped.jpg", "first_frame"),
                     frame("https://example.com/candle.jpg", "last_frame")],
}
r = requests.post(f"{API}/v1/videos", json=body,
                  headers={**H, "Idempotency-Key": "unwrap-001"})
r.raise_for_status()
job = r.json()
for _ in range(60):
    s = requests.get(job["polling_url"], headers=H).json()
    if s["status"] in ("completed", "failed"):
        break
    time.sleep(15)
print(s["status"], s.get("unsigned_urls"), s.get("error"))

What to check

Check the middle, not just the ends. Sume's Video frames route extracts source-size stills at times you name, so ask for 1, 2.5, 4 and 5.8 seconds.

  • The last still should be your product photo, with the same label and color.
  • The paper should not turn into a different object halfway: look for a box that changes size or a lid that vanishes.
  • If the model renders sound, a rustle is a bonus; do not promise it, because the sound is generated with the picture.
  • Both frames must be public HTTPS images in a supported format, or the request fails before generation.

Draft at the cheapest resolution that shows the motion, then re-render only the winner. No video model on Sume accepts a seed, so a final render is a new generation and may unwrap a little differently.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume