AI virtual staging video: from empty room to furnished

An AI virtual staging video stages the empty-room photo first, then animates it or reveals the furniture, with the empty photo first and the staged one last.

5 min readSume
All posts

An AI virtual staging video shows an empty room furnished, as a short clip. It is made in two steps: first stage a still, an AI image edit of the empty-room photo, then turn that still into video. Either animate the staged room with a slow camera move, or make a reveal that starts on the real empty room and ends on the staged one.

The Sume steps below come from the Image API, Video generation, Media inputs and Timeline 1.0 docs, read on 2026-09-29. Limits marked as current behavior are read from Sume's code.

How do I stage the still first?

Send the empty-room photo to POST /v1/images as an input_references entry, with a prompt naming the furniture style and what must not change; on edit calls the docs recommend aspect_ratio: "auto" to match the reference. Virtual staging AI covers the prompt, the models and the output size.

The result's data[].url is Sume-hosted and signed. Video inputs must be fetchable public HTTPS URLs, and signed or private URLs are rejected, so copy the staged still you keep to a public HTTPS location of your own before the next step.

How do I make the empty-to-staged reveal?

Send both photos in frame_images on one POST /v1/videos job: the real empty room as first_frame and the staged still as last_frame. The model generates the frames in between. Pick a model whose supported_frame_images on GET /v1/videos/models lists last_frame; in current code a last_frame without a first_frame is refused. For a plain walk-through of the staged room instead, send only the staged still as first_frame with a slow camera move in the prompt.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: living-room-reveal-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "Static camera. Furniture fades into the empty room. Walls, windows and floor do not change.",
    "duration": 5,
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/empty.jpg" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/staged.jpg" }, "frame_type": "last_frame" }
    ]
  }'

Will the room stay accurate in the video?

Only the frames you send are fixed. Everything in between is generated, so walls, windows, floors and the view can shift, and every piece of furniture is invented. Watch each clip before you use it, and cut any that changes the room itself.

Listing sites and local rules differ on whether and how virtually staged media must be labeled. Check the rules of each site where the video will appear; this post is not legal advice.

How do I put several staged rooms into one tour?

Join the clips in room order with one Timeline 1.0 render over a voice-over or music. It reads only this workspace's media.sume.com files, such as the clips Sume returned. Real estate photo to video AI walks through the tour.

What does an AI virtual staging video cost?

You pay per step from your workspace USD balance. One room is one image edit and one video clip:

From the Image API, Video generation, Timeline 1.0 and the API pricing rate card, read 2026-09-29. Rates are plus a 5.5% agent fee by default.
StepCallPrice
Stage the still (image edit)POST /v1/imagesBy model: the pricing lines on GET /v1/images/models; a failed generation is not billed
Animate or revealPOST /v1/videosBy model, at provider list × 1.25 per clip; see pricing_skus on GET /v1/videos/models
Join several rooms (optional)POST /v1/timeline-1.0/render$0.10 per output minute

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume