AI architecture video generator: renders to moving shots

An AI architecture video generator animates a still render: the render is the first frame, the prompt names one camera move. Limits, walkthroughs, cost.

5 min readSume
All posts

An AI architecture video generator turns a still render or photo of a building into a short moving shot: the render becomes the clip's first frame, and the prompt names one camera move, such as a slow dolly in, a partial orbit or a crane up. Clips are short, so a walkthrough is several shots joined in an edit. The result is generated video, not a render of your model: everything beyond the frames you pin is invented, so it can't stand in for a walkthrough that has to match the plans.

The Sume steps below come from the Video generation and Timeline 1.0 docs, read on 2026-09-29. Limits marked as current behavior are read from Sume's code.

How do I turn an architectural render into a video?

Put the render at a public HTTPS URL and send it in frame_images with frame_type first_frame on POST /v1/videos. That makes the job image-to-video, so the shot opens on your exact image. The request has no camera setting, so the move goes in the prompt, one move per shot; AI video camera movement prompts has the vocabulary.

  • To end the move on a second render, say a closer view of the entrance, add it as a last_frame entry. In current code last_frame without first_frame is refused.
  • Send resolution explicitly, and read each model's supported_frame_images before relying on a last frame.
  • Wait for completed on the polling URL, then download from unsigned_urls with your API key.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facade-dolly-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "Slow dolly in toward the entrance of the building at dusk. The building does not change; only the camera moves. No people, no cars.",
    "duration": 6,
    "resolution": "1080p",
    "aspect_ratio": "16:9",
    "generate_audio": false,
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/renders/facade.png" }, "frame_type": "first_frame" }
    ]
  }'

How long can an AI architecture video be?

Each clip is capped by its model, so check GET /v1/videos/models before planning shots:

From Video generation, read 2026-09-29.
SettingWhat the docs say
Length per clipseedance-2.5 4–30 s, wan-3.0 2–30 s; the others at most 15 s (supported_durations per model)
Aspect ratiosInclude 16:9, 9:16, 1:1 and ultra-wide 21:9; each model lists its subset in supported_aspect_ratios
ResolutionsFrom 480p to 4K; each model lists its subset in supported_resolutions
Soundgenerate_audio defaults to the model's audio capability; send false for a silent shot

Can AI make a full architectural walkthrough?

Only as a sequence of short generated shots, joined in order: approach, entrance, lobby, a room, the view. The joining step is the same as for a listing tour from photos: the generated clips go in order into one Timeline 1.0 render (POST /v1/timeline-1.0/render) of up to 200 slots and 1,800 seconds. What changes for design renders:

  • Each shot starts from its own render, so consecutive shots don't share a continuous camera path. To carry one move into the next shot, see chaining clips from a last frame.
  • In current code the render keeps only the audio spine and the soundtrack, so any sound a clip generated is dropped. Put narration or music on the spine instead.
  • The default output is 1080×1920, vertical. For a 16:9 presentation, set output.width 1920 and output.height 1080.

Is an AI architecture video accurate enough to show clients?

Treat it as mood, not documentation. Only the frames you pin are your images; the model fills in everything between them.

  • Anything the render doesn't show, such as a side facade or the room behind a door, is invented. A long orbit shows more of it.
  • Geometry, materials, window counts and signage can drift between frames. Watch every clip before it reaches a client.
  • Don't rely on it for dimensions, finishes or anything a buyer or planner might treat as a commitment. For off-plan marketing, label the footage as illustrative.
  • A walkthrough that must match the drawings still comes from your 3D tool; an AI clip can be the moving teaser around it.

How much does an AI architecture video cost?

You pay per step from your workspace USD balance: one video job per shot, an optional narration, and one render per cut.

From Video generation, Timeline 1.0 and the API pricing rate card, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
StepCallPrice
Animate each renderPOST /v1/videosBy model, reserved at provider list × 1.25 per clip; see pricing_skus on GET /v1/videos/models
Narration (optional)POST /v1/tts-1.0/generate$0.0475 per 1,000 characters
Join the shots into one MP4POST /v1/timeline-1.0/render$0.10 per output minute

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume