Which AI video models can't do text-to-video? Rows that need a source

Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.

4 min readSume
All posts

Three rows in Sume's video catalog cannot generate from a prompt alone: Grok Imagine Video 1.5 needs a first-frame image, Genjutsu Motion Transfer needs one source video plus 1 to 8 reference images, and H3 Max Recast needs one source video plus 1 to 4 photos. Every other model, including Seedance 2.5, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash, accepts a text-only request.

What each row needs

Send the missing input and the request validates; leave it out and it fails with a 400 before any provider work starts.

Rows without text-to-video, read 2026-10-03. Source: Sume video catalog on origin/main, read 2026-10-03.
Model idRequired inputOutput
grok-imagine-video-1.5image_url (first frame)4 to 15 s, 480p or 720p, no audio
higgsfield-genjutsuvideo_url plus 1 to 8 reference_image_urlssource length and framing kept, 480p or 720p
h3-max-recastvideo_url plus 1 to 4 reference_image_urlssource length, motion and soundtrack kept, 768p or 1080p

Picking between them

Grok Imagine is the row for animating a still you already like. Genjutsu moves your reference images with the motion of a source video. Recast swaps the people in a source video for the people in your photos, one photo per person. These are three different jobs, not three prices for the same job.

Genjutsu is listed only when its provider is configured, so it may not appear in GET /v1/video-router/models in every environment. Check the live list.

If you only have a prompt

Pick a text-to-video row. Seedance 2.5 and Wan 3.0 reach 30 seconds, Kling 3 stops at 15, and Gemini Omni Flash 1.1 gives 3 to 10 seconds with native audio. Or generate a still first with an image model and use it as a first frame.

Sources

Related posts

More in Models

All Models posts

Written by Sume