Floor plan to video AI: animate a plan, and what to check
Floor plan to video AI means image-to-video: the plan is the first frame and the prompt sets the camera. The clip is an illustration, not a survey.

To turn a floor plan into a video with AI, use image-to-video: send the plan image as the clip's first frame and describe the camera move in the prompt, such as a slow top-down push across the rooms. The result is a moving illustration of the plan. It is not a measured 3D model, so don't use it to show room sizes.
The Sume facts below come from Video generation and Timeline 1.0, read 2026-09-29. Sume's docs don't describe any model that reads dimensions or walls from a plan, so this post makes no such claim.
How do I animate a floor plan?
Send POST /v1/videos with frame_images and one entry whose frame_type is first_frame. The docs describe frame_images as image-to-video and say it takes precedence when you also send input_references, so the plan itself is the opening frame.
Ask for one simple move per clip and name the rooms you want the camera to reach. GET /v1/videos/models lists which models accept first_frame and their durations; most models top out at 15 seconds, while seedance-2.5 and wan-3.0 accept longer clips.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "Slow top-down push-in over the floor plan, soft light spreading room by room, no text changes",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 8,
"frame_images": [{
"type": "image_url",
"image_url": { "url": "https://example.com/floor-plan.png" },
"frame_type": "first_frame"
}]
}'First frame or reference image: which should a plan use?
The two image fields do different jobs, per the docs. Pick by what you need to survive in the clip.
| Field | What the docs say | Use for a plan |
|---|---|---|
frame_images | First or last frame images for image-to-video | When the clip should open on the plan itself |
input_references | Style or content references; visual guidance rather than exact frames | When you want the look of a plan or a render, not the plan as frame 1 |
Will the video keep my room layout?
Treat it as not guaranteed. The image fixes the opening frame; every frame after it is generated. Text, labels and dimensions in a plan can drift, and walls or doors can move. Watch the whole clip against the source plan before it goes into a listing, and keep a note that it is an illustration.
For a room-by-room tour, generate one short clip per area and join them. Timeline 1.0 joins clips that are already media.sume.com files, in order, with transitions like fade or dissolve of at most 1 second on slots after the first. The existing post on a listing tour from photos covers that step.
What does it cost?
Video is billed by model, resolution and length, and the poll response reports usage.cost for the finished job. Read live prices from GET /v1/videos/models (pricing_skus) and the API pricing page rather than from a post. The docs also note that higher resolutions take longer and cost more, so pick the lowest one that reads clearly.
Sources
Related posts
More in Use cases
- Gemini Omni in YouTube Create: countries, length and limits
Gemini Omni in YouTube Create makes vertical clips up to 10 seconds on mobile in five countries. Google's limits, and how to get a 9:16 clip from Sume instead.
- Sketch to image with GPT Image 2.5 API: send a drawing as a reference
To turn a sketch into a finished image with GPT Image 2.5 on Sume, send the drawing in input_references and describe the result. The request, size tips, limits.
- Gym promo video with AI: from your own gym photos
Make a gym promo video without a shoot: animate photos of your own floor and classes, add a voiced offer, music and captions, and render a vertical cut.
- Hair salon promo video with AI, from your own photos
Make a hair salon promo video from your own photos: animate the space and real finished looks, add a voiced offer and music, and render it vertical.
Written by Sume