How to write a prompt for an AI video generator
Write an AI video prompt as one shot: subject and action, camera, light, and framing. A template, worked examples, and what goes in settings instead.

To write a prompt for an AI video generator, describe one shot in plain sentences: who or what is on screen, where, and what they do, then the camera's angle and movement, the light, and the composition. Put the length, shape, and resolution in the request's settings rather than the prompt, and describe what you want to see rather than what you don't.
Those are the details Sume's Video generation docs ask for: “details about motion, camera angles, lighting, and scene composition.” The request fields and the short example prompts below come from that page, read on 2026-09-28. The template and the longer examples are illustrations, not wording any model is documented to follow.
What should an AI video prompt include?
One shot per prompt, in this order. Each line of the template below adds one part to the same example:
- Subject and action: who or what, doing what. “A golden retriever chases a tennis ball along the waterline.”
- Setting: where and when. “A sunny beach, waves breaking behind it.”
- Camera: angle and movement. “Low tracking shot at the dog's eye level, moving left to right.”
- Light: source, time of day, and mood. “Late-afternoon sun, warm light, long shadows.”
- Composition: framing, and what stays where. “The dog stays centered; the horizon sits in the upper third.”
What are some AI video prompt examples?
Sume's docs use one-line prompts in their request examples: “A golden retriever playing fetch on a sunny beach with waves crashing in the background”, “A character walking through a forest”, “A colossal solar flare beside a planet”, “A time-lapse of a flower blooming”, and “A vertical UGC-style product clip on a desk, natural light”. Three of them, expanded with the template:
- “A golden retriever runs along the waterline of a sunny beach, chasing a tennis ball, waves crashing behind it. Low tracking shot at the dog's eye level. Warm late-afternoon light. The dog stays centered.”
- “A woman in a green raincoat walks along a forest path toward the camera. The camera moves slowly backward ahead of her at chest height. Soft morning light through the trees, light mist. She stays in the center of the frame.”
- “A flower bud opens into full bloom in a time-lapse. Static macro close-up from the side. Soft, even light on a dark background. The flower fills the center of the frame.”
What goes in the settings instead of the prompt?
Anything the request has a field for. The fields are fixed: in current code, a field the API doesn't define is refused with a 400 before any job starts.
| Choice | Where it goes | Notes |
|---|---|---|
| Action, camera, light, composition | prompt | The text description of the video |
| Length | duration | Whole seconds from the model's supported_durations |
| Resolution | resolution | A value from supported_resolutions, such as 720p |
| Shape | aspect_ratio | A value from supported_aspect_ratios, such as 16:9 or 9:16 |
| Sound | generate_audio | Defaults to the model's audio capability |
| Opening image | frame_images | A first_frame, and optionally a last_frame |
| Look | input_references | Reference images for style guidance |
| Things to leave out | No field | Describe what you want instead; see AI video negative prompt |
Should the prompt describe the photo I start from?
It doesn't need to. With image-to-video, the photo is the clip's first frame, so the prompt is better spent on what changes after it. Image to video prompt examples covers that case, and Camera movement prompts for AI video covers the camera part of the template.
How do I improve a prompt that isn't working?
Change one part of the template at a time and keep every setting the same between runs, so you can see what the change did. No video model on Sume accepts a seed, so each run is a new take rather than a replay; What is an AI video seed? covers what to hold fixed instead. For resolution words like “4K” in a prompt, see AI video 4K quality prompt.
Sources
Related posts
More in Models
- Image to image AI: turn your photo into a new picture
Image-to-image AI redraws your photo from a text instruction: a new style, background or color. What it changes, what it keeps, and how to run it.
- Is Kling AI Chinese? Who owns it and where it's based
Yes. Kling AI is developed by Kuaishou Technology, a Beijing-based company listed in Hong Kong. Who owns it, where its API runs, and other access.
- AI video prompt JSON: what the keys actually do
A JSON prompt for AI video is still text: the model reads keys as words. Length, size and audio are request fields, not keys inside the prompt.
- Kling 3.0 API: text or image to video, limits and price
Kling 3.0 has an API: Kling's own, and multi-model APIs such as Sume's, where it is kling-3: 4–15 second clips from text or frames, audio optional.
Written by Sume