Lyria 3.5 takes ten images; the Sume music router takes one
Google documents up to 10 images with text for Lyria 3.5. The Sume music router accepts a single image_url, so combine your mood board into one still first.

Google's Lyria 3.5 documentation says the model accepts up to 10 images together with text, but the Sume Music Router takes one optional image_url per request. If you have a mood board, flatten it into one composite still and send that.
This is a Sume limit, not a Lyria one, so state it that way when someone asks why a ten-image workflow from Google's docs does not carry over directly.
What Google documents
The Gemini API music generation page describes Lyria 3 Clip as a fixed 30-second MP3 and Lyria 3.5 as running a couple of minutes, with MP3 or WAV output at 44.1 kHz stereo. For 3.5 it allows up to 10 images alongside the text prompt, supports custom lyrics with section tags such as [Verse] and [Chorus], and adds a SynthID watermark. It also notes that lyrics follow the prompt language and that requests to imitate an artist's voice or reproduce copyrighted lyrics are blocked.
| Model | Length | Output | Images |
|---|---|---|---|
| lyria-3-clip-preview | Always 30 seconds | MP3 | See the doc |
| lyria-3.5 | A couple of minutes | MP3 or WAV, 44.1 kHz stereo | Up to 10 with text |
What Sume accepts
The Music Router docs list sume/music-auto (currently Lyria 3.5), lyria-3.5 and lyria-3-pro. The prompt is 1 to 5,000 characters, and one optional image_url can accompany it. Duration fields are rejected and negative_prompt is not supported. Output is an audio artifact, and each generation has a fixed price listed in the docs. The retiring Music 1.0 resolves through the router.
That is the whole contract: one prompt, one image at most. If you send several image references, the request does not match the schema.
Flatten a mood board
Place three to six reference frames on one canvas, with generous spacing, and export a single JPEG or PNG. The model will see one picture, so avoid text overlays and keep colors representative of the mood you want. Describe in the prompt what the image should contribute, for example palette, setting or energy, so the picture supports the words instead of competing with them.
- Keep the image publicly reachable at the URL you send.
- Say in text what to take from the image.
- Specify genre, tempo feel and instrumentation in the prompt, since duration cannot be set.
- Re-run with a tightened prompt rather than adding more images.
When ten images would matter
If you need Google's full multi-image behavior, that is available through Google's own API, outside Sume. If a single composite and a good prompt get you to a usable bed, the Sume route keeps the audio in the same job history as your video work. The image-to-music post shows a worked example.
Sources
Related posts
More in Models
- MAI-Voice-2.1-Flash makes 45 s of audio: narrating a 5-minute script
Microsoft lists 45 seconds of audio per Flash generation. A 5-minute script needs about seven chunks. A Python splitter, and how Sume handles long text.
- MAI-Voice-2.1 on OpenRouter and Vercel: is it in Sume's TTS Router?
Microsoft lists MAI voices on Foundry, OpenRouter and Vercel. Sume's TTS Router lists only Cartesia Sonic ids today. How to check the live catalog.
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
- MiniMax H3 or H3 Max on Sume: which id for a 5 to 15 second clip?
Choose minimax-h3 for a 480p or 768p draft at $0.075 a second, minimax-h3-max when you need 1080p. Same 5-15 s range, same references; Max costs more at 768p.
Written by Sume