Home renovation before and after video: start and end frame

Make a renovation reveal clip from two photos: send the before as image_url and the after as end_image_url to Gemini Omni Flash 1.1, 3 to 10 seconds.

5 min readSume
All posts

To make a before and after video of a home renovation with AI, send the before photo as image_url and the after photo as end_image_url to gemini-omni-flash-1.1 on the Video Router. The model animates the move from one frame to the other in 3 to 10 seconds, at 360p up to 4K, in 16:9 or 9:16, and Sume's docs say it always generates synced audio.

Both photos are yours, so the first and last frames are real; only the seconds between them are generated. That is the honest framing for a contractor: the clip is a transition, not a record of the work. Every field below is from the Video Router page and the Video generation guide, read on 2026-10-04.

What the request looks like

Gemini Omni Flash 1.1 is one catalog id that Sume routes by the shape of the request. A request with image_url is image-to-video, and adding end_image_url makes it a start-and-end-frame request on the same envelope. There is no endpoint to pick.

Gemini Omni Flash 1.1 image-to-video fields (read 2026-10-04)
FieldValueNote
modelgemini-omni-flash-1.1Routed by request shape
image_urlThe before photoPublic HTTPS URL
end_image_urlThe after photoOptional second frame
duration3 to 10 secondsWhole seconds
resolution360p, 720p, 1080p or 4KBilling is provider list x 1.25 per output second by resolution
aspect_ratio16:9 or 9:16Match the shape of your two photos
generate_audioAlways onSending false is rejected

Submit it

Shoot the after photo from the same spot and height as the before photo, because the model has to bridge two frames and a changed camera position turns the bridge into a cut. Keep the prompt about the camera and the light, not the new finishes, since the finishes are already in the last frame.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: reno-reveal-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Locked-off camera in a kitchen. Dust sheets lift away and the room changes from the first frame to the last frame. Natural daylight, no people.",
    "image_url": "https://example.com/kitchen-before.jpg",
    "end_image_url": "https://example.com/kitchen-after.jpg",
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "duration": 8,
    "mode": "async"
  }'

Check the ends, drop the audio if you want your own

Because native audio is always on, you cannot ask for a silent clip. If the reveal should sit under a music track instead, cut the clip with Video trim and audio: "drop", which is a flat $0.02 job and returns a new MP4 (not the source). Then lay your own bed under it in Timeline.

Before you post, pull the first and last frames with Video frames and compare them with your two photos. The at values must fall inside the clip, so for an 8-second clip ask for [0, 7.5].

  • Frame at 0 should match the before photo; if the start drifted, regenerate rather than trim.
  • Frame at 7.5 should match the after photo's layout and furniture.
  • If the middle shows rooms or fixtures that were never in either photo, say the transition is AI-made in the caption.

When not to use it

Gemini Omni's edit mode is a different request: video_url is an edit source and cannot be combined with image_url or end_image_url. If you already have a real time-lapse and only want to change one thing in it, use the edit mode; if you only have two stills, use this one.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume