Image-to-video prompt: write the motion, not the photo
When a photo is already pinned as the first frame, spend the prompt on motion, camera and sound, not on describing what the picture shows. A worked rewrite.

When your photo is pinned as the first frame, write the prompt about what changes, not about what is already visible. The image already tells the model the subject, the setting and the light. A prompt that repeats it ("a woman in a red coat on a street at dusk") spends your words on information the model has, and says nothing about the part only you can supply: what moves, how the camera moves, and what the clip sounds like.
This is practice advice, not a Sume rule. What Sume does document is the mechanism: an image in frame_images with first_frame conditions the opening frame, and the video guide asks for prompts with detail about motion, camera angle, lighting and scene composition. Put those two together and the prompt for an image-to-video request is mostly verbs.
A prompt in three parts
A short structure keeps a motion prompt from turning into a caption. Write one sentence for each part and stop. Longer is not better when the picture already carries the scene.
- Subject motion: what the subject does in the next seconds, such as turns her head, lifts the cup, walks toward the camera.
- Camera: static, slow push-in, pan left, handheld sway. Name one move.
- Atmosphere that moves: steam, rain, leaves, traffic in the background, and any sound the model makes natively.
Before and after
Here is the same opening frame with a descriptive prompt and with a motion prompt. The first reads like alt text. The second is something a model can act on. Both are plain text sent in the prompt field with the photo in frame_images.
| Version | Prompt | Problem or gain |
|---|---|---|
| Descriptive | A barista in a green apron holds a latte in a bright cafe | Repeats the photo, no motion asked for |
| Motion | She slides the latte across the counter, slow push-in, steam curls up, soft cafe chatter | Subject, camera and moving atmosphere |
| Over-written | Cinematic masterpiece, 8K, ultra detailed, perfect hands, no blur | Quality adjectives, no motion |
What to leave to the request fields
Resolution, aspect ratio and duration are fields, not prompt words. Send resolution and duration in the body and check the model's limits first, because they differ: Seedance 2.5 takes 4 to 30 seconds, Wan 3.0 takes 2 to 30 seconds, and Gemini Omni Flash 1.1 takes 3 to 10 seconds. Writing "ten seconds" in the prompt does not change the duration that the API bills.
Native sound is handled the same way. Gemini Omni Flash 1.1, MiniMax H3 and MiniMax H3 Max always generate audio, so there is nothing to switch on, and a prompt can describe the sound you want. Models with an audio toggle follow the generate_audio field.
Three common mistakes
The first mistake is describing the photo again. The second is asking for a change that contradicts the pinned frame, for example a night scene on a photo taken at noon, which forces the model to fight the first frame it was given. If you need a change of lighting, use a first and last frame pair, or edit the source image before you animate it. The third is stacking five motions in one sentence. A clip of a few seconds can hold one clear action and one camera move, and more than that tends to read as noise.
Length helps explain the third mistake. Wan 3.0 starts at 2 seconds and Gemini Omni Flash 1.1 at 3 seconds, so a short clip has room for exactly one idea. If the story needs more beats, plan more clips and join them on a timeline instead of packing the beats into one prompt.
Check and iterate
Sume has no seed on video models, so two runs of the same prompt give two different clips. Treat the prompt as a brief you iterate: change one thing, run again, and compare the results. If the clip starts on the wrong frame, the cause is probably the field and not the wording, which is covered in why a clip does not start on your photo. For negative wording, see how to write exclusions for video.
Sources
Related posts
More in Models
- Nano Banana 2 is retired on Sume: same price on Nano Banana 2.1?
Sume retired google/nano-banana-2 and runs old requests as Nano Banana 2.1 at the same price card: $0.10 at 1K billed. Tiers, ids and what the job stores.
- Remove an object from a photo without a mask: Ideogram 4.5 prompt edit
Remove a bin, cable or stray person from a photo with no mask: send it to ideogram/ideogram-v4.5 and name the object. $0.0375 to $0.275 per image on Sume.
- Replace one prop in a video with AI: bottle to apple edit prompt
Swap one object in a finished clip with gemini-omni-flash-1.1 on Sume: the docs example prompt, fields you cannot set, and an 8-second price.
- Square 1:1 ad video: which Sume models accept it, and 4:5
Seedance, Kling, Wan and MiniMax accept 1:1 on Sume; Gemini Omni takes only 16:9 and 9:16. No video model lists 4:5, so Feed needs a crop. Table and steps.
Written by Sume