How does AI video generation work? Inputs, jobs, and limits
AI video generation turns a prompt, and optionally an image, into every frame of a short clip in one async job. What goes in, and why clips stay short.

AI video generation works like this: you give a video model a text prompt, and optionally a starting image or reference images, and the model generates every frame of a short clip, and on some models its sound, in one job. Because that typically takes 30 seconds to several minutes, video APIs run it as an asynchronous job that you submit, check on, and download.
The mechanics below are described with Sume's Video generation API, which serves several video model families behind one POST /v1/videos request, plus the Models and Timeline 1.0 docs, read on 2026-09-28. Model internals differ by vendor and are not covered here.
What goes into an AI video generator?
A text prompt and the model id are the only required inputs. Everything else is either a picture the model starts from or a setting for the output. If you send both frame images and references, the frame images win and the request is treated as image-to-video.
| Input | What it does |
|---|---|
prompt | Text description of the video (required) |
frame_images | A first_frame image, optionally with a last_frame: image-to-video. In current code a last_frame without a first_frame is refused |
input_references | Style or content references the model uses as visual guidance, not as exact frames |
duration | Clip length in seconds |
resolution | Output resolution, such as 720p or 1080p |
aspect_ratio | Frame shape, such as 16:9 or 9:16 |
generate_audio | Whether to make an audio track with the video; defaults to the model's audio capability |
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "A paper boat drifting down a rain-soaked street, low angle",
"duration": 5,
"resolution": "720p",
"aspect_ratio": "9:16"
}'Why does AI video generation take so long?
Generating video takes much longer than text or images, so the API does not hold the connection open. The docs say video generation typically takes 30 seconds to several minutes, depending on the model and parameters. The flow is:
- Submit the request. The answer is
202 Acceptedwith a jobid, apolling_url, andstatus: pending. - Poll
GET /v1/videos/{jobId}until the status iscompleted, or send an HTTPScallback_urlto get a webhook instead. - Download the file from
GET /v1/videos/{jobId}/contentwith the same API key. It redirects to the file, so curl needs-L.
Why are AI videos so short?
Each model caps the length of one request. On Sume, seedance-2.5 accepts 4–30 seconds and wan-3.0 2–30 seconds; every other catalog model tops out at 15 seconds. AI video length limits by model lists every range.
Length and size also drive cost, because the model has to produce more frames and more pixels. In current code, for example, Sume's Seedance cost estimate counts video tokens as (width * height * duration_seconds * 24_fps) / 1024, and the docs note that higher resolutions take longer to generate and cost more. Why are AI videos so expensive? breaks the prices down.
How are longer AI videos made?
By generating several clips and joining them. On Sume the joining step is Timeline 1.0, which renders one audio spine and ordered video slots into one MP4 at $0.10 per output minute; How to assemble a long-form video walks through a render. Every URL must already be a Sume-hosted file, and in current code the render takes sound only from the spine and an optional soundtrack, so each clip's own audio is dropped.
A talking person is a separate step too. Sume's docs say video models do not lip-sync to generated speech or a later voice-over, so a speaking face is made with a lip-sync model from a still and an audio file; see Lip sync API.
What does a video model not do?
For a hands-on walkthrough of prompting and picking an input, see How to make AI videos.
- It does not make a finished long video in one request; each clip stays within its model's length range.
- It does not lip-sync a voice-over you add later.
- It supports only some resolutions, aspect ratios, and durations; check them with
GET /v1/videos/modelsbefore you submit.
Sources
Related posts
More in Models
- How to download Sora videos after the shutdown
To download Sora videos, open sora.chatgpt.com/sunset and click Export; OpenAI emails you when it's ready. Export soon: the data is deleted later.
- How to write a prompt for an AI video generator
Write an AI video prompt as one shot: subject and action, camera, light, and framing. A template, worked examples, and what goes in settings instead.
- Image to image AI: turn your photo into a new picture
Image-to-image AI redraws your photo from a text instruction: a new style, background or color. What it changes, what it keeps, and how to run it.
- Image to video prompt examples: what to write
An image-to-video prompt needn't describe the photo again. It says what moves, what the camera does, and what stays still. Examples by photo type.
Written by Sume