Text-to-video or image-to-video: when the prompt alone is enough

Use text-to-video when the look is open, and image-to-video when a frame, a face or a product must match. The Sume models that take each, and a decision rule.

5 min readSume
All posts

Use text-to-video when you only need a scene that matches a description and any good version will do. Use image-to-video when something in the first frame has to be exact: a face, a product, a logo, a layout or a drawn style. If you would reject a result because one detail is wrong, you need an image, because a prompt cannot promise an identical detail twice.

A common mistake is to use image-to-video for everything out of habit. When the content is generic, an image adds a step and a constraint with no gain, and a poor starting image makes a poor clip. The other mistake is to use text only for something that has to match, then spend an afternoon re-rolling to get close. The question is not which mode is better but which detail you cannot afford to lose.

On Sume both are the same endpoint, POST /v1/videos. The mode is inferred from the body: a body with only a prompt is text-to-video, and a body with frame_images is image-to-video. So the choice costs you a field, not a different integration.

The decision rule in four questions

Answer these before the first request. If every answer is no, the prompt alone is enough.

  • Must a specific person, product or place appear? If yes, use an image.
  • Must the clip begin on a particular composition? If yes, use a first frame.
  • Must the clip match a style you already have? If yes, use an image of that style.
  • Do you need to compare several models on the same start? If yes, use one image for all of them.

What text-to-video is good at

Text-to-video works best for scenes that are generic by design: weather, landscapes, crowds, abstract motion, establishing shots and b-roll. It is also the cheapest way to explore, because you do not have to make an image first. Write the subject, the action, the camera and the light in one or two sentences, and run a few variations to find a direction.

Treat each text-only clip as a sketch. It tells you whether an idea works in motion, and that is valuable, but it is not a stable asset. Expect to throw most of them away, and plan the budget on that basis: run short, low resolution clips and keep only the frame or two that you want to build on.

It is a poor tool for continuity. Two text-to-video clips from the same prompt show different people, different rooms and different light, so a sequence made this way does not hold together. When the second clip has to look like the first, move to images.

What image-to-video adds

An image fixes the starting frame, and the prompt then only has to describe motion. That is a shorter and easier prompt, covered in writing the motion, not the photo. The result starts on your frame, which is what you cannot get from words. If you need the clip to be conditioned by an image without pinning the first frame, that is reference-to-video and a different field, explained in image-to-video versus reference-to-video.

Which Sume models take each

The capabilities below come from the Video Router catalog. Read GET /v1/videos/models for the live version.

Input modes by model (Sume docs, read 2026-10-07)
Model idText-to-videoImage-to-videoReference-to-video
wan-3.0YesYes, with optional end frameYes
minimax-h3YesYes, with optional end frameYes
minimax-h3-maxYesYes, with optional end frameYes
gemini-omni-flash-1.1YesYes, with optional end frameYes
kling-3YesYesNo reference inputs
grok-imagine-video-1.5NoImage-to-video onlyNo

A practical workflow

Keep a record of the prompt that produced the frame you liked, together with the model and the duration. When you move to image-to-video, shorten the prompt to the motion alone, since the image now carries the subject, the clothing and the light. A prompt that repeats what the image already shows can pull the result away from the image.

Start with text-to-video at the lowest resolution to find a direction. When you like a frame, save it, then use it as the first frame for the real run. This keeps the exploring cheap and gives the final clip a start you chose.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume