How to make AI videos from text or a photo

To make an AI video, pick what it starts from, describe the shot, set its length and shape, then generate and review. How to do each step on Sume.

5 min readSume
All posts

To make an AI video, decide what the clip starts from (words only, or a photo as its first frame), write a prompt that describes the shot, choose a length, shape, and resolution the model supports, then generate the clip and review it. Each generation is a short clip, so a longer video is several clips joined together.

On Sume, one clip runs 2 to 30 seconds depending on the model, and you can do all of this in a chat with the agent in the Agents tab, or by sending the same choices to the video API. Facts come from Sume's Quick start, Video generation, and Timeline 1.0 docs, read on 2026-09-28.

Can I make an AI video from text or from my own photo?

Both. The fields you send decide the mode, and if you send frames and references together, the frames win. Image to video vs reference to video explains the difference.

From Video generation and Media inputs, read 2026-09-28.
Start fromWhat you sendWhat the model does with it
Words only (text-to-video)A promptGenerates the shot from your description
Your photo (image-to-video)The photo's public HTTPS URL as first_frame in frame_images, plus a promptUses it as the clip's first frame; a last_frame can set the end
Reference images (reference-to-video)input_references, plus a promptUses them as visual guidance, not exact frames

What should the prompt say?

Describe one shot: what is on screen and what happens, the camera, the light, and the framing. Sume's docs ask for “details about motion, camera angles, lighting, and scene composition”, and their own example reads “A golden retriever playing fetch on a sunny beach with waves crashing in the background.” Length, shape, and resolution are request settings, not prompt words. How to write a prompt for an AI video generator has a template.

How long can one AI video be?

If you let Sume pick the model (sume/auto), a clip defaults to 720p and 8 seconds, and runs 3–10 seconds at 16:9 or 9:16. Pinning a model changes the range: seedance-2.5 and wan-3.0 go to 30 seconds, and the rest stop at 15 or less. AI video length limits by model lists them.

For a longer video, generate several clips and join them, for example in one Timeline 1.0 render. Its sound comes only from the audio track you give it and an optional soundtrack; in current code, each clip's own sound is dropped.

Can I make an AI video without code?

Yes. Agents is Sume's chat UI: you write a brief, approve spend, and inspect what it made. Sume's quick start puts it this way: “Describe the deliverable, not the tool calls. The agent picks the models, and asks before it spends.” What is a video agent? covers what it can build beyond one clip.

How do I make an AI video with the API?

Four steps, per the docs: submit to POST /v1/videos, receive a job id and a polling URL, poll GET /v1/videos/{jobId} until the status is completed, then download from GET /v1/videos/{jobId}/content with your API key. sume/auto lets Sume pick the model and never says which one ran. Sume API quickstart walks through the first call.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: first-video-001" \
  -d '{
    "model": "sume/auto",
    "prompt": "A vertical UGC-style product clip on a desk, natural light",
    "aspect_ratio": "9:16",
    "duration": 5
  }'

How long does it take, and what does it cost?

The docs say video generation typically takes 30 seconds to several minutes, depending on the model and parameters, and that higher resolutions take longer and cost more. Each clip is priced at the provider's list price × 1.25, reserved from your workspace balance when you submit, plus a 5.5% agent fee by default. How much does an AI video cost? has the rates by model.

Sources

Related posts

More in Models

All Models posts

Written by Sume