Earth zoom out AI video: from your photo to space

Make an Earth zoom out AI video by pinning both ends: your photo as the first frame, the Earth from space as the last, and one zoom in the prompt.

5 min readSume
All posts

To make an Earth zoom out video with AI, give an image-to-video model both ends of the move: your close-up photo as the first frame, a picture of the whole Earth from space as the last frame, and a prompt that describes one continuous pull-back between them. For an Earth zoom in, swap the two frames. The model generates everything in between, so the streets, coastlines, and clouds it passes are invented, not real map imagery.

Facts come from Sume's Video generation, Image API, and Media inputs docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.

How do I make an Earth zoom out video?

Without code, describe the zoom in the Agents tab: the agent picks the models and asks before it spends. Over the API:

  • Pick the start photo: a person, a rooftop, or a landmark, cropped to the shape you will ask for, such as 9:16 for a vertical clip.
  • Get the end image in the same shape: a photo of Earth from space that you may use, or a still you generate from a text prompt with POST /v1/images.
  • Host both images at public HTTPS URLs. The Image API returns signed result URLs, and signed or private URLs are refused as generation inputs.
  • Send POST /v1/videos with frame_images: your photo as first_frame and the Earth as last_frame, on a model that takes a last frame.
  • Poll the job until its status is completed, then watch the whole move before you post it.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: earth-zoom-out-001" \
  -d '{
    "model": "seedance-2.5",
    "prompt": "One continuous zoom out: the camera rises straight up from the woman on the rooftop, over the city, the coastline and the clouds, until the whole Earth hangs in black space",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/rooftop.jpg" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/earth.jpg" }, "frame_type": "last_frame" }
    ],
    "aspect_ratio": "9:16",
    "resolution": "720p",
    "duration": 12
  }'

What should the Earth zoom out prompt say?

The request has no camera field, so the whole move goes in prompt. Sume's docs suggest details about motion, camera angles, lighting, and scene composition, and camera movement prompts covers the general vocabulary. For this effect that means:

  • One move in one direction, with no cuts: "one continuous zoom out" or "the camera rises straight up".
  • The layers it passes, in order: the street, the rooftops, the city, the coastline, the clouds, the curve of the Earth.
  • An ending that matches your last frame: "until the whole Earth fills the frame against black space".
  • For a zoom in, reverse both the words and the frames, as in the table.
Frame fields from Video generation, read 2026-09-28. The prompts are examples, not settings.
Effect`first_frame``last_frame`Prompt, for example
Earth zoom outYour photoThe Earth from space"One continuous zoom out: the camera rises from the woman on the rooftop, over the city and the clouds, until the whole Earth fills the frame"
Earth zoom inThe Earth from spaceYour photo"One continuous zoom in: the camera falls from space through the clouds toward the city and lands on the man on the bench"
Zoom with stopsOne stop's stillThe next stop's stillOne leg per clip: "the camera rises from the street until the whole city fills the frame"

How long can the zoom be?

As long as one clip: up to 30 seconds on seedance-2.5 or wan-3.0, and 15 seconds or less on the other models, as AI video length limits by model lists. A longer clip gives the move more time to pass each layer. Every model except grok-imagine-video-1.5 takes the last frame that the Earth image needs.

Can the zoom stop at more places on the way out?

Yes, by chaining clips. Make one still per stop, such as the street, the city from above, the continent, and the Earth, then generate one clip per neighboring pair so that each clip ends on the image the next one starts from; AI video from multiple images covers chaining pairs. Join the clips in one Timeline 1.0 render, which reads only your workspace's media.sume.com files, as generated clips already are. The render takes sound only from its audio spine and optional soundtrack: in current code each clip's own audio is dropped.

What are the limits?

  • The frames between your two images are generated. Nothing promises a smooth zoom or real geography, so treat the middle of the clip as illustration.
  • A last_frame needs a first_frame in current code, and grok-imagine-video-1.5 takes no last frame.
  • Frame images must be at public HTTPS URLs. Localhost, private-network, and signed or private URLs are refused.
  • Each clip is billed by model; How much does an AI video cost? has the rates.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume