Text to video and image to video: how they differ
Text-to-video invents every frame from words; image-to-video starts on your picture and animates it. How they differ, and how one request does both.

Text-to-video generates a clip from words alone, so the model invents every frame, including the first. Image-to-video starts the clip on a picture you supply and uses the prompt to describe what moves and changes after it. On Sume's video API they are the same request: add a first-frame image to a text-to-video request and it becomes image-to-video.
The request rules come from Video generation and the Sume API reference, and per-model support from the catalog behind GET /v1/videos/models, read on 2026-09-28; "current code" marks behavior read from the API code. A third mode, reference-to-video, uses pictures as guidance without pinning a frame; Image to video vs reference to video covers it.
How does image-to-video differ from text-to-video?
| Item | Text-to-video | Image-to-video |
|---|---|---|
| You send | prompt | prompt plus frame_images with a first_frame |
| First frame | Invented by the model | Your picture |
| Last frame | Invented by the model | Invented, or a second picture sent as last_frame |
| Models on Sume | Every listed model except grok-imagine-video-1.5 | Every listed model; grok-imagine-video-1.5 takes a first frame only |
| Price | Set by the model and the clip's settings | The same rate: a frame doesn't change it |
Can one request use both text and an image?
Yes, and it always does: prompt is required on every request, and frame_images adds the picture. Each entry has a frame_type of first_frame or last_frame; in current code a last_frame without a first_frame is refused. The image must be a public HTTPS URL. This is the docs' image-to-video example; drop frame_images and it is a text-to-video request.
{
"model": "seedance-2",
"prompt": "A character walking through a forest",
"frame_images": [
{
"type": "image_url",
"image_url": { "url": "https://example.com/first-frame.png" },
"frame_type": "first_frame"
}
],
"resolution": "1080p"
}Which should I use, text or image to video?
- Text-to-video when you have no picture, or any scene that matches the words will do.
- Image-to-video when the clip must open on a specific product, person, place or design: the model starts from your picture instead of inventing one.
- Add a
last_framewhen the ending must be fixed too, such as a before-and-after or a transition. First and last frame covers the per-model rules. - With an image, the prompt only has to say what moves; Image to video prompt examples shows how to write it.
Does image-to-video cost more than text-to-video?
No. In current code a clip's price comes from the model and the clip's settings: its length on every model, its resolution on most, its aspect ratio on Seedance (which counts video tokens), and whether audio is on for Kling. A first or last frame is not part of the price, and pricing_skus has no mode key: rates are per second, per second by resolution, or per 1,000 video tokens.
Which models do only one of the two?
One: grok-imagine-video-1.5 is image-to-video only, and a request without a first frame is refused. The others do both: seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, minimax-h3, minimax-h3-max, gemini-omni-flash-1.1. With a picture, crop it to the shape you want: in current code kling-3, minimax-h3 and minimax-h3-max don't pass aspect_ratio to the model when a first frame is sent. Grok Imagine video API covers the image-only model.
Sources
Related posts
More in Developers
- Webhook security best practices: a receiver checklist
Accept HTTPS only, verify an HMAC over the raw body in constant time, reject stale timestamps, dedupe on the event id, and answer 2xx fast.
- Webhook vs API: what's the difference?
An API call is your code asking a server for something; a webhook is the server calling your URL when something happens. The two work together.
- What is a dead letter queue? DLQs for AI job pipelines
A dead-letter queue holds messages that failed processing too many times, so they stop looping and can be inspected. Which AI job failures go there.
- What is a video API? The five kinds, explained
A video API lets code make, edit or deliver video over HTTP. The five kinds, what each one takes and returns, and how to tell which one you need.
Written by Sume