Text-to-video or image-to-video: when the prompt alone is enough
Use text-to-video when the look is open, and image-to-video when a frame, a face or a product must match. The Sume models that take each, and a decision rule.

Use text-to-video when you only need a scene that matches a description and any good version will do. Use image-to-video when something in the first frame has to be exact: a face, a product, a logo, a layout or a drawn style. If you would reject a result because one detail is wrong, you need an image, because a prompt cannot promise an identical detail twice.
A common mistake is to use image-to-video for everything out of habit. When the content is generic, an image adds a step and a constraint with no gain, and a poor starting image makes a poor clip. The other mistake is to use text only for something that has to match, then spend an afternoon re-rolling to get close. The question is not which mode is better but which detail you cannot afford to lose.
On Sume both are the same endpoint, POST /v1/videos. The mode is inferred from the body: a body with only a prompt is text-to-video, and a body with frame_images is image-to-video. So the choice costs you a field, not a different integration.
The decision rule in four questions
Answer these before the first request. If every answer is no, the prompt alone is enough.
- Must a specific person, product or place appear? If yes, use an image.
- Must the clip begin on a particular composition? If yes, use a first frame.
- Must the clip match a style you already have? If yes, use an image of that style.
- Do you need to compare several models on the same start? If yes, use one image for all of them.
What text-to-video is good at
Text-to-video works best for scenes that are generic by design: weather, landscapes, crowds, abstract motion, establishing shots and b-roll. It is also the cheapest way to explore, because you do not have to make an image first. Write the subject, the action, the camera and the light in one or two sentences, and run a few variations to find a direction.
Treat each text-only clip as a sketch. It tells you whether an idea works in motion, and that is valuable, but it is not a stable asset. Expect to throw most of them away, and plan the budget on that basis: run short, low resolution clips and keep only the frame or two that you want to build on.
It is a poor tool for continuity. Two text-to-video clips from the same prompt show different people, different rooms and different light, so a sequence made this way does not hold together. When the second clip has to look like the first, move to images.
What image-to-video adds
An image fixes the starting frame, and the prompt then only has to describe motion. That is a shorter and easier prompt, covered in writing the motion, not the photo. The result starts on your frame, which is what you cannot get from words. If you need the clip to be conditioned by an image without pinning the first frame, that is reference-to-video and a different field, explained in image-to-video versus reference-to-video.
Which Sume models take each
The capabilities below come from the Video Router catalog. Read GET /v1/videos/models for the live version.
| Model id | Text-to-video | Image-to-video | Reference-to-video |
|---|---|---|---|
| wan-3.0 | Yes | Yes, with optional end frame | Yes |
| minimax-h3 | Yes | Yes, with optional end frame | Yes |
| minimax-h3-max | Yes | Yes, with optional end frame | Yes |
| gemini-omni-flash-1.1 | Yes | Yes, with optional end frame | Yes |
| kling-3 | Yes | Yes | No reference inputs |
| grok-imagine-video-1.5 | No | Image-to-video only | No |
A practical workflow
Keep a record of the prompt that produced the frame you liked, together with the model and the duration. When you move to image-to-video, shorten the prompt to the motion alone, since the image now carries the subject, the clothing and the light. A prompt that repeats what the image already shows can pull the result away from the image.
Start with text-to-video at the lowest resolution to find a direction. When you like a frame, save it, then use it as the first frame for the real run. This keeps the exploring cheap and gives the final clip a start you chose.
Sources
Related posts
More in Comparisons
- Can I use Suno Speech beta for an ad voiceover? What to check first
Suno Speech beta makes one track with voice and music. Before using it for ads, check price, languages and edits, then see how Sume splits voice and music.
- Zoom in on a small product in frame: AI recompose or crop and upscale
Product too small in the photo? A Pillow crop plus Sume Image Upscale ($0.20) keeps real pixels; an Ideogram 4.5 recompose edit costs $0.075 and redraws them.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume