Veo 3.1 Lite takes text and image only; Sume models take more inputs
Google lists Veo 3.1 Lite inputs as text and image. On Sume, Seedance 2.0 and Wan 3.0 also accept video and audio references; Kling 3 does not.

The Gemini API's Veo 3.1 Lite page lists input as text and image and output as video with audio, one video per request, with no 4K and no Extension (read 2026-10-04). If your brief includes a reference clip or a voice track, that input list rules it out.
On Sume, the supported_input_references field of each catalog row tells you which of image_url, video_url and audio_url a model accepts, per the video generation docs.
Inputs by row
Wan 3.0 video and audio references are capped at 15 seconds in total each. Omni video references are each 3 seconds or shorter.
| Model | Text | Image | Video ref | Audio ref |
|---|---|---|---|---|
| Veo 3.1 Lite (Google, not on Sume) | Yes | Yes | No | No |
| seedance-2 | Yes | Yes | Yes | Yes |
| wan-3.0 | Yes | Yes (up to 10) | Yes (up to 5) | Yes (up to 5) |
| gemini-omni-flash-1.1 | Yes | Yes (up to 10) | Yes (up to 3) | No |
| kling-3 | Yes | Start and end frame only | No | No |
Pick by the input you have
Check the live row before building, since catalogs change.
- Only a prompt and a product photo: any row works, and Kling is a good fit for start-frame motion.
- A reference clip for motion: Seedance 2.0 or Wan 3.0.
- A voice or music track to drive the clip: Seedance 2.0 or Wan 3.0, since Omni takes no audio reference.
Sources
Related posts
More in Models
- Video aspect ratios by Sume row: Kling 3, Gemini, Wan 3.0, MiniMax H3
Which aspect_ratio values each Sume video row takes: kling-3 has three, Gemini two, Wan and MiniMax add adaptive, and MiniMax adds 21:9. From the catalog.
- MAI-Voice-2.1-Flash tops out at 45 seconds: what for longer?
MAI-Voice-2.1-Flash is reported at up to 45 s of audio and 150 ms latency. Longer lines need the standard model or a TTS taking 20,000 characters.
- Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.
- Wan 3.0 doubles clip length to 30 seconds: Alibaba ids vs Sume
Alibaba Model Studio lists wan3.0-video and wan3.0-video-prime at 2 to 30 seconds, up from 15 on Wan 2.7. What Sume's wan-3.0 accepts and a 30 s request.
Written by Sume