Lyria 3.5 takes up to 10 images; Sume music takes one image_url
Google's Lyria guide allows up to 10 images as input. Sume's music docs describe one optional public HTTPS image_url. What that changes.

Google's Lyria guide says you can give a prompt up to 10 images. Sume's Music 1.0 docs describe one optional image_url, a public HTTPS link. Read 2026-10-01. If you want more than one picture to set the mood, describe the rest in words.
How do I use one image?
Pass image_url with your prompt, for example a cover or a moodboard frame, and say in the text what the image means for the music.
| Service | Image input |
|---|---|
| Google Lyria guide | Up to 10 images |
| Sume Music 1.0 | One optional public HTTPS image_url |
What if I have a moodboard of ten?
Pick the most telling frame, or write the mood, palette and setting into the prompt (up to 5,000 characters). Sume's docs make no promise about how closely music follows a picture.
Does the image replace the text prompt?
No. Keep the text prompt (up to 5,000 characters) as the main instruction and use the image as extra context.
Sources
Related posts
More in Models
- MAI-Transcribe-2 languages (60) vs Sume STT language hint
Microsoft's MAI-Transcribe-2 preview covers 60 languages. Sume STT 1.0 takes one optional language_code hint and auto-detects when you omit it.
- MAI-Transcribe-2 speaker diarization vs Sume STT: no speaker field
MAI-Transcribe-2 attributes each segment to a distinct speaker. Sume STT has no diarize field: provider knobs are fixed server-side and you get words[].
- Midjourney edit model image references: 4 vs Sume's ranges
Midjourney's V8.2 edit model takes up to 4 image references. On Sume the ceiling is per model: GPT Image 2.5 takes up to 16. Read the descriptor first.
- MiniMax-H3 Fun ControlNet Union 2.0: 8 conditions vs Sume references
MiniMax-H3-Fun-Controlnet-Union-2.0 adds Scribble, Layout and Gray to five older conditions. Sume's minimax-h3 ids take image, video and audio references.
Written by Sume