Image-to-video prompt: write the motion, not the photo

When a photo is already pinned as the first frame, spend the prompt on motion, camera and sound, not on describing what the picture shows. A worked rewrite.

5 min readSume
All posts

When your photo is pinned as the first frame, write the prompt about what changes, not about what is already visible. The image already tells the model the subject, the setting and the light. A prompt that repeats it ("a woman in a red coat on a street at dusk") spends your words on information the model has, and says nothing about the part only you can supply: what moves, how the camera moves, and what the clip sounds like.

This is practice advice, not a Sume rule. What Sume does document is the mechanism: an image in frame_images with first_frame conditions the opening frame, and the video guide asks for prompts with detail about motion, camera angle, lighting and scene composition. Put those two together and the prompt for an image-to-video request is mostly verbs.

A prompt in three parts

A short structure keeps a motion prompt from turning into a caption. Write one sentence for each part and stop. Longer is not better when the picture already carries the scene.

  • Subject motion: what the subject does in the next seconds, such as turns her head, lifts the cup, walks toward the camera.
  • Camera: static, slow push-in, pan left, handheld sway. Name one move.
  • Atmosphere that moves: steam, rain, leaves, traffic in the background, and any sound the model makes natively.

Before and after

Here is the same opening frame with a descriptive prompt and with a motion prompt. The first reads like alt text. The second is something a model can act on. Both are plain text sent in the prompt field with the photo in frame_images.

Rewrite pattern for a pinned first frame (practice example, read 2026-10-07 against Sume docs)
VersionPromptProblem or gain
DescriptiveA barista in a green apron holds a latte in a bright cafeRepeats the photo, no motion asked for
MotionShe slides the latte across the counter, slow push-in, steam curls up, soft cafe chatterSubject, camera and moving atmosphere
Over-writtenCinematic masterpiece, 8K, ultra detailed, perfect hands, no blurQuality adjectives, no motion

What to leave to the request fields

Resolution, aspect ratio and duration are fields, not prompt words. Send resolution and duration in the body and check the model's limits first, because they differ: Seedance 2.5 takes 4 to 30 seconds, Wan 3.0 takes 2 to 30 seconds, and Gemini Omni Flash 1.1 takes 3 to 10 seconds. Writing "ten seconds" in the prompt does not change the duration that the API bills.

Native sound is handled the same way. Gemini Omni Flash 1.1, MiniMax H3 and MiniMax H3 Max always generate audio, so there is nothing to switch on, and a prompt can describe the sound you want. Models with an audio toggle follow the generate_audio field.

Three common mistakes

The first mistake is describing the photo again. The second is asking for a change that contradicts the pinned frame, for example a night scene on a photo taken at noon, which forces the model to fight the first frame it was given. If you need a change of lighting, use a first and last frame pair, or edit the source image before you animate it. The third is stacking five motions in one sentence. A clip of a few seconds can hold one clear action and one camera move, and more than that tends to read as noise.

Length helps explain the third mistake. Wan 3.0 starts at 2 seconds and Gemini Omni Flash 1.1 at 3 seconds, so a short clip has room for exactly one idea. If the story needs more beats, plan more clips and join them on a timeline instead of packing the beats into one prompt.

Check and iterate

Sume has no seed on video models, so two runs of the same prompt give two different clips. Treat the prompt as a brief you iterate: change one thing, run again, and compare the results. If the clip starts on the wrong frame, the cause is probably the field and not the wording, which is covered in why a clip does not start on your photo. For negative wording, see how to write exclusions for video.

Sources

Related posts

More in Models

All Models posts

Written by Sume