AI music from an image: Lyria 3.5 image_url mood still on Sume
Sume's Music Router takes an optional public HTTPS image_url as a mood still next to the prompt. How to send it, what it costs and what to expect back.

What image_url does
The Music Router request has a required prompt of 1 to 5,000 characters and an optional image_url: a public HTTPS image, or null to clear. It acts as a mood still that accompanies the text. The default model is sume/music-auto, which routes to Lyria 3.5, and lyria-3.5 and lyria-3-pro can be named.
fal's Lyria 3.5 page lists the price as $0.1 per generation and describes MP3 output with lyrics carrying section markers.
A request
Send it to POST /v1/music-router/generate with an Idempotency-Key. There is no duration field, so say the length you want in the prompt.
{
"prompt": "Warm lo-fi piano with soft rain, about 30 seconds",
"image_url": "https://example.com/mood.jpg",
"model": "sume/music-auto"
}What it costs and returns
The Sume price is $0.125, the fal list price of $0.10 plus the Sume margin.
| Item | Value |
|---|---|
| Price | $0.125 per generation |
| Prompt | 1 to 5,000 characters |
| Image | Optional public HTTPS image_url |
| Output | Audio artifact (audio/mpeg) in result.artifacts[] |
| Lyrics | result.lyrics, a section map or lyrics when the model reports them |
| Engine used | job.request.routed_model |
Getting better results
- Pick a still with a clear mood: a rainy window, a neon street, a sunlit kitchen.
- Keep the prompt concrete about instruments and tempo, since the image sets mood, not structure.
- Do not send
negative_prompt; a non-empty value returnsnegative_prompt_unsupported. - Send
durationand you get a 400, so state length in words and trim with a split afterward.
Making a family of tracks
Because there is no seed field in the request, treat each generation as a fresh take. To build a family of related tracks, keep the same image and a stable stem in the prompt, and vary one thing at a time: tempo, lead instrument, or length. Tag each with metadata so you can find the keeper.
At $0.125 per generation, ten variations cost $1.25.
Where it fits
Use it when you have a product or scene image and want a bed that matches it, such as a still from a video you are about to cut. Generate, listen, and keep the one that fits. For fixed-length edits, see exact length music. Pricing is on the API pricing page.
Sources
Related posts
More in Media tools
- AI music generator for covers: what Sume cannot do, and a swap
Sume cannot cover an existing song: no audio input and no song-to-song path. Licensed cover platforms are still in development. Here is an original-track swap.
- Apple Podcasts transcript speaker names: VTT, and Sume STT gaps
Apple shows speaker names when you provide a VTT. Sume STT has no speaker labels, so transcribe each track and merge with voice tags. Script included.
- Audio-only cut of a 5 minute episode: mp3 for $0.01
Turn a finished 5 minute video into a 128 kbps mp3 for $0.01 with audio-detach. Source up to 1800 s, output up to 900 s, and the silent-source error.
- Batch trim clips from a spreadsheet of start and end times
Read start and end columns from a CSV and trim a long video into Shorts: one Video trim call per row at $0.02, with idempotency keys.
Written by Sume