Image to video prompt examples: what to write
An image-to-video prompt needn't describe the photo again. It says what moves, what the camera does, and what stays still. Examples by photo type.

An image-to-video prompt doesn't need to describe the picture again, because the photo is already the clip's first frame. Spend the words on what changes after it: the motion, the camera, and what must stay still. One image-to-video example in Sume's docs is a single line: “Gentle camera drift; keep the product locked in frame.”
That example comes from Sume's Video 1.0 page, which is retiring; new requests go to POST /v1/videos, whose own image-to-video example is just as short: “A character walking through a forest.” Video 1.0 and Image 1.0 are retiring soon covers the move. The other facts come from the Video generation docs and Sume's guidance for its own agent, which lives in its code, read on 2026-09-28.
How is an image-to-video prompt different?
In text-to-video, the prompt has to describe everything. In image-to-video, your photo fixes the opening frame: Sume's guidance for its own agent says “frame_images pins frame 0”. The prompt only has to cover what happens next, and a second image can fix where the clip ends.
| Input | What it does |
|---|---|
first_frame in frame_images | Your photo becomes the clip's first frame |
last_frame in frame_images | Sets the clip's last frame; in current code it needs a first_frame |
input_references | Visual guidance, not exact frames; when frames are sent too, the frames take precedence |
prompt | What happens after the first frame: motion, camera, light |
What are some image-to-video prompt examples?
Each one assumes the photo is sent as the first frame. They are illustrations to adapt, not wording a model is documented to follow:
- Portrait: “She slowly turns her head toward the camera and smiles; a light breeze moves her hair. The camera stays still.”
- Product: “Gentle camera drift; keep the product locked in frame.” That is the docs' example; Product video prompt examples has more.
- Landscape: “Clouds drift across the sky and the lake ripples in the wind. Slow pan from left to right.”
- Pet: “The cat blinks, then stretches and yawns. The camera stays at its eye level.”
- Full-body motion: “The whole person walks toward the camera, full body in frame, legs visible.” For body-motion clips, Sume's guidance for its own agent is to prompt full-body and whole-person, with legs in frame.
- Two frames: with a wide shot of a room as the first frame and a close-up of its armchair as the last, describe the move between them: “The camera slowly pushes in from the full room to the armchair by the window.”
What should an image-to-video prompt leave out?
- A description of what the photo already shows. It is the first frame.
- A different ending scene in words. If the clip should end on a particular picture, send it as the
last_frame. - Length, shape, and resolution. Those are the request fields
duration,aspect_ratio, andresolution. - A style to copy from other images. Style references go in
input_references, but with a first frame in the same request the frames take precedence, and in current code the references are dropped.
How do I send the photo and the prompt?
Put the photo at a public HTTPS URL and send it as the first_frame to POST /v1/videos. Signed links are refused, so a still from the Image API, which returns signed URLs, has to be hosted publicly first. How to make a picture move with AI covers choosing the model.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: portrait-motion-001" \
-d '{
"model": "seedance-2",
"prompt": "She slowly turns her head toward the camera and smiles; a light breeze moves her hair. The camera stays still.",
"frame_images": [
{
"type": "image_url",
"image_url": { "url": "https://example.com/portrait.png" },
"frame_type": "first_frame"
}
],
"resolution": "720p"
}'Sources
Related posts
More in Models
- Is Kling AI Chinese? Who owns it and where it's based
Yes. Kling AI is developed by Kuaishou Technology, a Beijing-based company listed in Hong Kong. Who owns it, where its API runs, and other access.
- AI video prompt JSON: what the keys actually do
A JSON prompt for AI video is still text: the model reads keys as words. Length, size and audio are request fields, not keys inside the prompt.
- Kling 3.0 API: text or image to video, limits and price
Kling 3.0 has an API: Kling's own, and multi-model APIs such as Sume's, where it is kling-3: 4–15 second clips from text or frames, audio optional.
- Kling 3.0 prompt guide: length, shots, dialogue, negatives
A Kling 3.0 prompt can run to 3,072 characters and include negative wording. Kling's rules for shot lists, dialogue, languages and Elements.
Written by Sume