How to make AI videos from text or a photo
To make an AI video, pick what it starts from, describe the shot, set its length and shape, then generate and review. How to do each step on Sume.

To make an AI video, decide what the clip starts from (words only, or a photo as its first frame), write a prompt that describes the shot, choose a length, shape, and resolution the model supports, then generate the clip and review it. Each generation is a short clip, so a longer video is several clips joined together.
On Sume, one clip runs 2 to 30 seconds depending on the model, and you can do all of this in a chat with the agent in the Agents tab, or by sending the same choices to the video API. Facts come from Sume's Quick start, Video generation, and Timeline 1.0 docs, read on 2026-09-28.
Can I make an AI video from text or from my own photo?
Both. The fields you send decide the mode, and if you send frames and references together, the frames win. Image to video vs reference to video explains the difference.
| Start from | What you send | What the model does with it |
|---|---|---|
| Words only (text-to-video) | A prompt | Generates the shot from your description |
| Your photo (image-to-video) | The photo's public HTTPS URL as first_frame in frame_images, plus a prompt | Uses it as the clip's first frame; a last_frame can set the end |
| Reference images (reference-to-video) | input_references, plus a prompt | Uses them as visual guidance, not exact frames |
What should the prompt say?
Describe one shot: what is on screen and what happens, the camera, the light, and the framing. Sume's docs ask for “details about motion, camera angles, lighting, and scene composition”, and their own example reads “A golden retriever playing fetch on a sunny beach with waves crashing in the background.” Length, shape, and resolution are request settings, not prompt words. How to write a prompt for an AI video generator has a template.
How long can one AI video be?
If you let Sume pick the model (sume/auto), a clip defaults to 720p and 8 seconds, and runs 3–10 seconds at 16:9 or 9:16. Pinning a model changes the range: seedance-2.5 and wan-3.0 go to 30 seconds, and the rest stop at 15 or less. AI video length limits by model lists them.
For a longer video, generate several clips and join them, for example in one Timeline 1.0 render. Its sound comes only from the audio track you give it and an optional soundtrack; in current code, each clip's own sound is dropped.
Can I make an AI video without code?
Yes. Agents is Sume's chat UI: you write a brief, approve spend, and inspect what it made. Sume's quick start puts it this way: “Describe the deliverable, not the tool calls. The agent picks the models, and asks before it spends.” What is a video agent? covers what it can build beyond one clip.
How do I make an AI video with the API?
Four steps, per the docs: submit to POST /v1/videos, receive a job id and a polling URL, poll GET /v1/videos/{jobId} until the status is completed, then download from GET /v1/videos/{jobId}/content with your API key. sume/auto lets Sume pick the model and never says which one ran. Sume API quickstart walks through the first call.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: first-video-001" \
-d '{
"model": "sume/auto",
"prompt": "A vertical UGC-style product clip on a desk, natural light",
"aspect_ratio": "9:16",
"duration": 5
}'How long does it take, and what does it cost?
The docs say video generation typically takes 30 seconds to several minutes, depending on the model and parameters, and that higher resolutions take longer and cost more. Each clip is priced at the provider's list price × 1.25, reserved from your workspace balance when you submit, plus a 5.5% agent fee by default. How much does an AI video cost? has the rates by model.
Sources
Related posts
More in Models
- How to write a prompt for an AI video generator
Write an AI video prompt as one shot: subject and action, camera, light, and framing. A template, worked examples, and what goes in settings instead.
- Image to image AI: turn your photo into a new picture
Image-to-image AI redraws your photo from a text instruction: a new style, background or color. What it changes, what it keeps, and how to run it.
- Image to video prompt examples: what to write
An image-to-video prompt needn't describe the photo again. It says what moves, what the camera does, and what stays still. Examples by photo type.
- Is Kling AI Chinese? Who owns it and where it's based
Yes. Kling AI is developed by Kuaishou Technology, a Beijing-based company listed in Hong Kong. Who owns it, where its API runs, and other access.
Written by Sume