Veo 3.1 Lite takes text and image only; Sume models take more inputs

Google lists Veo 3.1 Lite inputs as text and image. On Sume, Seedance 2.0 and Wan 3.0 also accept video and audio references; Kling 3 does not.

3 min readSume
All posts

The Gemini API's Veo 3.1 Lite page lists input as text and image and output as video with audio, one video per request, with no 4K and no Extension (read 2026-10-04). If your brief includes a reference clip or a voice track, that input list rules it out.

On Sume, the supported_input_references field of each catalog row tells you which of image_url, video_url and audio_url a model accepts, per the video generation docs.

Inputs by row

Wan 3.0 video and audio references are capped at 15 seconds in total each. Omni video references are each 3 seconds or shorter.

Reference inputs from the Sume catalog; Veo 3.1 Lite from the Google page read 2026-10-04
ModelTextImageVideo refAudio ref
Veo 3.1 Lite (Google, not on Sume)YesYesNoNo
seedance-2YesYesYesYes
wan-3.0YesYes (up to 10)Yes (up to 5)Yes (up to 5)
gemini-omni-flash-1.1YesYes (up to 10)Yes (up to 3)No
kling-3YesStart and end frame onlyNoNo

Pick by the input you have

Check the live row before building, since catalogs change.

  • Only a prompt and a product photo: any row works, and Kling is a good fit for start-frame motion.
  • A reference clip for motion: Seedance 2.0 or Wan 3.0.
  • A voice or music track to drive the clip: Seedance 2.0 or Wan 3.0, since Omni takes no audio reference.

Sources

Related posts

More in Models

All Models posts

Written by Sume