Image to drone shot AI: turn a photo into an aerial move
To turn an image into a drone shot, send it as the first frame of an AI video and describe the aerial move. Everything past the photo is invented.

To turn an image into a drone shot with AI, send the photo to an image-to-video model as the clip's first frame and describe an aerial camera move in the prompt, such as rising over the scene, flying forward, or orbiting a subject. The model generates the flight from that frame on, so whatever lies beyond the photo's edges is invented. The result is a generated shot, not footage of the real place.
Facts come from Sume's Video generation, Video Router, and Media inputs docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's API code.
How do I turn a photo into a drone shot?
Without code, describe the shot in the Agents tab: the agent picks the models and asks before it spends. Over the API:
- Pick a photo with room to fly: open sky, a skyline, a coastline, or a building seen from above. Crop it to the shape you want, such as 16:9.
- Put it at a public HTTPS URL. Localhost, private-network, non-HTTPS, and signed or private URLs are refused.
- Send it to
POST /v1/videosinframe_imagesas thefirst_frame, and write the move inprompt. - Set an
aspect_ratioand adurationthe model lists, and setresolutionexplicitly.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: drone-coast-001" \
-d '{
"model": "seedance-2",
"prompt": "Aerial drone shot: the camera rises slowly from the beach and flies forward over the cliffs toward the lighthouse, golden hour light",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/coast.jpg" }, "frame_type": "first_frame" }
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 8
}'What should an AI drone shot prompt say?
The request has no camera field, so the move goes in the prompt. Sume's docs suggest details about motion, camera angles, lighting, and scene composition, and camera movement prompts covers the general vocabulary. Name one move, its direction, and its speed:
- Rise and reveal: "The camera rises slowly from the rooftop and tilts down, revealing the whole city at dusk."
- Fly forward: "A low aerial shot flying over the waves toward the lighthouse."
- Orbit: "The camera circles the house at roof height, keeping it in the center of the frame."
- Pull back: "The camera flies backward and up until the whole island fits in the frame."
Can I make a drone shot without a photo?
Yes: send only a prompt that describes both the place and the move. Every model in the table below except grok-imagine-video-1.5 runs from text alone; that one needs a first frame. With a photo, you can also pin where the flight ends: add a last_frame next to the first_frame on a model that takes one, as in Earth zoom out AI video.
Will the drone shot come out in the shape I ask for?
It depends on the model. Each model lists the aspect_ratio values it accepts, and in current code a value it doesn't list is refused; Vertical 9:16 video generation API has the lists. With a first frame, though, current code passes aspect_ratio on to only some models, so crop the photo to the shape you want before you send it:
| Model id | `aspect_ratio` in a first-frame request |
|---|---|
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini | Checked, then sent to the model |
wan-3.0 | Checked, then sent to the model |
gemini-omni-flash-1.1 | Checked (16:9 or 9:16), then sent to the model |
kling-3 | Checked, but not sent to the model |
minimax-h3, minimax-h3-max | Checked, but not sent to the model |
grok-imagine-video-1.5 | Not accepted: the model lists none |
What are the limits?
- The surroundings are invented. A generated drone shot is not a record of a real property or piece of land; Real estate photo to video AI covers what that means for listings.
- The prompt steers the move, and no field sets it. Nothing promises the model follows it, so watch the clip.
- One clip runs 30 seconds at most, on
seedance-2.5orwan-3.0; AI video length limits by model lists the rest. For a longer flight, start the next clip on the last frame of the one before. - Each clip is billed by model, at the rates
GET /v1/videos/modelslists inpricing_skus.
Sources
Related posts
More in Use cases
- AI flyer generator: make the art, then type the details
Generate flyer art at US Letter or A4 shape with empty panels, then type the headline, date, prices, and address yourself so you can check each one.
- AI food photography: restyle your real dish or generate one
Photograph the real dish and let AI restyle the scene around it, or generate food from a prompt where no dish is promised. Prompts, steps, and limits.
- AI ghost mannequin: remove the mannequin, keep the shape
A ghost mannequin photo shows a garment's 3D shape with no mannequin. Make one with an AI image edit of your photos; cut it out if you need a PNG.
- AI greeting card generator: a card front from your photo
Make a greeting card with AI: restyle your own photo as the card's front at the card's shape, then write the message inside yourself.
Written by Sume