How does AI video generation work? Inputs, jobs, and limits

AI video generation turns a prompt, and optionally an image, into every frame of a short clip in one async job. What goes in, and why clips stay short.

5 min readSume
All posts

AI video generation works like this: you give a video model a text prompt, and optionally a starting image or reference images, and the model generates every frame of a short clip, and on some models its sound, in one job. Because that typically takes 30 seconds to several minutes, video APIs run it as an asynchronous job that you submit, check on, and download.

The mechanics below are described with Sume's Video generation API, which serves several video model families behind one POST /v1/videos request, plus the Models and Timeline 1.0 docs, read on 2026-09-28. Model internals differ by vendor and are not covered here.

What goes into an AI video generator?

A text prompt and the model id are the only required inputs. Everything else is either a picture the model starts from or a setting for the output. If you send both frame images and references, the frame images win and the request is treated as image-to-video.

From Video generation, read 2026-09-28.
InputWhat it does
promptText description of the video (required)
frame_imagesA first_frame image, optionally with a last_frame: image-to-video. In current code a last_frame without a first_frame is refused
input_referencesStyle or content references the model uses as visual guidance, not as exact frames
durationClip length in seconds
resolutionOutput resolution, such as 720p or 1080p
aspect_ratioFrame shape, such as 16:9 or 9:16
generate_audioWhether to make an audio track with the video; defaults to the model's audio capability
curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seedance-2",
    "prompt": "A paper boat drifting down a rain-soaked street, low angle",
    "duration": 5,
    "resolution": "720p",
    "aspect_ratio": "9:16"
  }'

Why does AI video generation take so long?

Generating video takes much longer than text or images, so the API does not hold the connection open. The docs say video generation typically takes 30 seconds to several minutes, depending on the model and parameters. The flow is:

  • Submit the request. The answer is 202 Accepted with a job id, a polling_url, and status: pending.
  • Poll GET /v1/videos/{jobId} until the status is completed, or send an HTTPS callback_url to get a webhook instead.
  • Download the file from GET /v1/videos/{jobId}/content with the same API key. It redirects to the file, so curl needs -L.

Why are AI videos so short?

Each model caps the length of one request. On Sume, seedance-2.5 accepts 4–30 seconds and wan-3.0 2–30 seconds; every other catalog model tops out at 15 seconds. AI video length limits by model lists every range.

Length and size also drive cost, because the model has to produce more frames and more pixels. In current code, for example, Sume's Seedance cost estimate counts video tokens as (width * height * duration_seconds * 24_fps) / 1024, and the docs note that higher resolutions take longer to generate and cost more. Why are AI videos so expensive? breaks the prices down.

How are longer AI videos made?

By generating several clips and joining them. On Sume the joining step is Timeline 1.0, which renders one audio spine and ordered video slots into one MP4 at $0.10 per output minute; How to assemble a long-form video walks through a render. Every URL must already be a Sume-hosted file, and in current code the render takes sound only from the spine and an optional soundtrack, so each clip's own audio is dropped.

A talking person is a separate step too. Sume's docs say video models do not lip-sync to generated speech or a later voice-over, so a speaking face is made with a lip-sync model from a still and an audio file; see Lip sync API.

What does a video model not do?

For a hands-on walkthrough of prompting and picking an input, see How to make AI videos.

  • It does not make a finished long video in one request; each clip stays within its model's length range.
  • It does not lip-sync a voice-over you add later.
  • It supports only some resolutions, aspect ratios, and durations; check them with GET /v1/videos/models before you submit.

Sources

Related posts

More in Models

All Models posts

Written by Sume