AI construction timelapse: from empty lot to finished build

An AI construction timelapse pins the empty site as the first frame and the finished building as the last. The model invents every stage between.

5 min readSume
All posts

An AI construction timelapse is a generated clip that starts on an empty site and ends on the finished building, with the stages in between invented by a video model. Send the site photo as the clip's first frame and an image of the finished building as its last frame, and ask for a time-lapse of the build in the prompt. The result illustrates a build; it is not a record of real work.

"Time-lapse" is prompt wording, not a setting: Sume's Video generation docs use "A time-lapse of a flower blooming" as an example prompt. The rest comes from those docs and the Image API, Media inputs, and Timeline 1.0 pages, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.

How do I make an AI construction timelapse?

  • Start from a fixed viewpoint: a photo of the empty lot or bare room, taken from where a time-lapse camera would stand.
  • Make the end frame from the same viewpoint. Use a render of the design, or edit the site photo on POST /v1/images: send it in input_references, describe the finished building in the prompt, and set aspect_ratio: "auto" to match the reference, on a model whose catalog lists it.
  • Host both images at public HTTPS URLs. The Image API returns signed result URLs, and signed or private URLs are refused as generation inputs.
  • Send POST /v1/videos with the site as first_frame and the building as last_frame, on a model that takes a last frame.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: build-timelapse-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A time-lapse of a two-story house being built on this lot: foundation, framing, roof, windows, then paint and landscaping. Fixed camera, clouds racing, day and night cycling",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/empty-lot.jpg" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/finished-house.jpg" }, "frame_type": "last_frame" }
    ],
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 15
  }'

What should an AI timelapse prompt say?

No request field turns on a time-lapse, so the prompt carries it. Sume's docs suggest details about motion, camera angles, lighting, and scene composition. For a build:

  • Say "time-lapse" and keep the camera still: "fixed camera" or "locked-off wide shot".
  • List the stages in the order they happen: foundation, framing, roof, windows, finish.
  • Show time passing in the light: "clouds racing", "day and night cycling", "shadows sweeping across the lot".
  • Keep one subject: the building. Workers and machines can pass through, but say what the frame is about.

How do I show more stages?

Pin them. Make one still per stage from the same viewpoint, such as the empty lot, the foundation, the frame, and the finished house, each as an edit of the one before. Generate one clip per neighboring pair, so that each clip ends on the still the next one starts from, then join the clips in one Timeline 1.0 render. AI video from multiple images covers chaining pairs.

From Image API, Video generation, and Timeline 1.0, read 2026-09-28. Clip lengths are from the catalog behind GET /v1/videos/models.
StepCallRules that matter
A still per stagePOST /v1/images with the previous still in input_referencesPublic HTTPS reference; prefer aspect_ratio: "auto" to match it; results come back at signed URLs, so host each still you keep publicly
A clip per pairPOST /v1/videos with frame_imagesA last_frame needs a first_frame; 4–30 s on seedance-2.5, 2–30 s on wan-3.0, 15 s or less elsewhere
One videoPOST /v1/timeline-1.0/renderYour workspace's media.sume.com clips only; 1–200 slots; audio.mode: "silence" when there is no sound

What are the limits?

  • The stages are invented. The clip is not progress footage or a record of the real build, and nothing promises the construction steps it shows are physically right.
  • A last_frame needs a first_frame in current code, and grok-imagine-video-1.5 takes no last frame.
  • One clip runs 30 seconds at most. For a longer timelapse, chain clips.
  • Timeline takes only Sume-hosted files. Generated clips qualify because Sume mirrors outputs to its own media URLs: once a job is completed, read the clip's URL from GET /v1/jobs/{id}/result.
  • Timeline's default output is 1080×1920, so set output.width and output.height for a 16:9 edit, such as 1920 and 1080.
  • A Timeline render takes sound only from its audio spine and optional soundtrack. In current code each clip's own audio is dropped.
  • For a reveal between two real photos of a finished project, see Before-and-after video generator API; for a listing, Real estate photo to video AI.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume