Lyria 3.5 takes ten images; the Sume music router takes one

Google documents up to 10 images with text for Lyria 3.5. The Sume music router accepts a single image_url, so combine your mood board into one still first.

4 min readSume
All posts

Google's Lyria 3.5 documentation says the model accepts up to 10 images together with text, but the Sume Music Router takes one optional image_url per request. If you have a mood board, flatten it into one composite still and send that.

This is a Sume limit, not a Lyria one, so state it that way when someone asks why a ten-image workflow from Google's docs does not carry over directly.

What Google documents

The Gemini API music generation page describes Lyria 3 Clip as a fixed 30-second MP3 and Lyria 3.5 as running a couple of minutes, with MP3 or WAV output at 44.1 kHz stereo. For 3.5 it allows up to 10 images alongside the text prompt, supports custom lyrics with section tags such as [Verse] and [Chorus], and adds a SynthID watermark. It also notes that lyrics follow the prompt language and that requests to imitate an artist's voice or reproduce copyrighted lyrics are blocked.

Lyria details from Google's documentation (read 2026-10-03)
ModelLengthOutputImages
lyria-3-clip-previewAlways 30 secondsMP3See the doc
lyria-3.5A couple of minutesMP3 or WAV, 44.1 kHz stereoUp to 10 with text

What Sume accepts

The Music Router docs list sume/music-auto (currently Lyria 3.5), lyria-3.5 and lyria-3-pro. The prompt is 1 to 5,000 characters, and one optional image_url can accompany it. Duration fields are rejected and negative_prompt is not supported. Output is an audio artifact, and each generation has a fixed price listed in the docs. The retiring Music 1.0 resolves through the router.

That is the whole contract: one prompt, one image at most. If you send several image references, the request does not match the schema.

Flatten a mood board

Place three to six reference frames on one canvas, with generous spacing, and export a single JPEG or PNG. The model will see one picture, so avoid text overlays and keep colors representative of the mood you want. Describe in the prompt what the image should contribute, for example palette, setting or energy, so the picture supports the words instead of competing with them.

  • Keep the image publicly reachable at the URL you send.
  • Say in text what to take from the image.
  • Specify genre, tempo feel and instrumentation in the prompt, since duration cannot be set.
  • Re-run with a tightened prompt rather than adding more images.

When ten images would matter

If you need Google's full multi-image behavior, that is available through Google's own API, outside Sume. If a single composite and a good prompt get you to a usable bed, the Sume route keeps the audio in the same job history as your video work. The image-to-music post shows a worked example.

Sources

Related posts

More in Models

All Models posts

Written by Sume