Text to video and image to video: how they differ

Text-to-video invents every frame from words; image-to-video starts on your picture and animates it. How they differ, and how one request does both.

4 min readSume
All posts

Text-to-video generates a clip from words alone, so the model invents every frame, including the first. Image-to-video starts the clip on a picture you supply and uses the prompt to describe what moves and changes after it. On Sume's video API they are the same request: add a first-frame image to a text-to-video request and it becomes image-to-video.

The request rules come from Video generation and the Sume API reference, and per-model support from the catalog behind GET /v1/videos/models, read on 2026-09-28; "current code" marks behavior read from the API code. A third mode, reference-to-video, uses pictures as guidance without pinning a frame; Image to video vs reference to video covers it.

How does image-to-video differ from text-to-video?

From Video generation and the catalog behind GET /v1/videos/models, read 2026-09-28.
ItemText-to-videoImage-to-video
You sendpromptprompt plus frame_images with a first_frame
First frameInvented by the modelYour picture
Last frameInvented by the modelInvented, or a second picture sent as last_frame
Models on SumeEvery listed model except grok-imagine-video-1.5Every listed model; grok-imagine-video-1.5 takes a first frame only
PriceSet by the model and the clip's settingsThe same rate: a frame doesn't change it

Can one request use both text and an image?

Yes, and it always does: prompt is required on every request, and frame_images adds the picture. Each entry has a frame_type of first_frame or last_frame; in current code a last_frame without a first_frame is refused. The image must be a public HTTPS URL. This is the docs' image-to-video example; drop frame_images and it is a text-to-video request.

{
  "model": "seedance-2",
  "prompt": "A character walking through a forest",
  "frame_images": [
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/first-frame.png" },
      "frame_type": "first_frame"
    }
  ],
  "resolution": "1080p"
}

Which should I use, text or image to video?

  • Text-to-video when you have no picture, or any scene that matches the words will do.
  • Image-to-video when the clip must open on a specific product, person, place or design: the model starts from your picture instead of inventing one.
  • Add a last_frame when the ending must be fixed too, such as a before-and-after or a transition. First and last frame covers the per-model rules.
  • With an image, the prompt only has to say what moves; Image to video prompt examples shows how to write it.

Does image-to-video cost more than text-to-video?

No. In current code a clip's price comes from the model and the clip's settings: its length on every model, its resolution on most, its aspect ratio on Seedance (which counts video tokens), and whether audio is on for Kling. A first or last frame is not part of the price, and pricing_skus has no mode key: rates are per second, per second by resolution, or per 1,000 video tokens.

Which models do only one of the two?

One: grok-imagine-video-1.5 is image-to-video only, and a request without a first frame is refused. The others do both: seedance-2.5, seedance-2-mini, seedance-2, seedance-2-fast, kling-3, wan-3.0, minimax-h3, minimax-h3-max, gemini-omni-flash-1.1. With a picture, crop it to the shape you want: in current code kling-3, minimax-h3 and minimax-h3-max don't pass aspect_ratio to the model when a first frame is sent. Grok Imagine video API covers the image-only model.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume