Which AI video models can't do text-to-video? Rows that need a source
Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.

Three rows in Sume's video catalog cannot generate from a prompt alone: Grok Imagine Video 1.5 needs a first-frame image, Genjutsu Motion Transfer needs one source video plus 1 to 8 reference images, and H3 Max Recast needs one source video plus 1 to 4 photos. Every other model, including Seedance 2.5, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash, accepts a text-only request.
What each row needs
Send the missing input and the request validates; leave it out and it fails with a 400 before any provider work starts.
| Model id | Required input | Output |
|---|---|---|
| grok-imagine-video-1.5 | image_url (first frame) | 4 to 15 s, 480p or 720p, no audio |
| higgsfield-genjutsu | video_url plus 1 to 8 reference_image_urls | source length and framing kept, 480p or 720p |
| h3-max-recast | video_url plus 1 to 4 reference_image_urls | source length, motion and soundtrack kept, 768p or 1080p |
Picking between them
Grok Imagine is the row for animating a still you already like. Genjutsu moves your reference images with the motion of a source video. Recast swaps the people in a source video for the people in your photos, one photo per person. These are three different jobs, not three prices for the same job.
Genjutsu is listed only when its provider is configured, so it may not appear in GET /v1/video-router/models in every environment. Check the live list.
If you only have a prompt
Pick a text-to-video row. Seedance 2.5 and Wan 3.0 reach 30 seconds, Kling 3 stops at 15, and Gemini Omni Flash 1.1 gives 3 to 10 seconds with native audio. Or generate a still first with an image model and use it as a first frame.
Sources
Related posts
More in Models
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Can you sell images from open-weights models? Licences compared
Open weights do not mean commercial use. Ideogram 4, Qwen-Image, FLUX.2 dev and LTX-2.5 differ on selling outputs. What each page says, and hosted rows.
- ChatGPT Try On from a screenshot: the same edit through an API
ChatGPT Try On starts from a selfie plus a product screenshot. Do the same edit with openai/gpt-image-2.5 on Sume: two references, one prompt, one Python call.
Written by Sume