Music from a thumbnail: image_url on Sume's music router
Generate a music bed that matches a still: pass one public HTTPS image_url with the prompt to Sume's Music Router, and clear it with null when reusing objects.
Can you generate music from an image? Yes, with Sume's Music Router: send a prompt and one optional image_url. The Music 1.0 docs call this visual conditioning and show a cinematic ambient underscore matched to the mood of a reference still. The image must be a public HTTPS URL.
An image does not replace the prompt. The prompt is still required (1 to 5000 characters), and the docs say to write exclusions and length into it.
The request
image_url is optional and can be null to clear. Send null only when you reuse a request object on a client and want the image gone.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: thumb-music-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Cinematic ambient underscore matching the mood of the reference still. Instrumental, no vocals. A 30-second track.",
"image_url": "https://example.com/moodboard.png"
}'Image conditioning, rule by rule
| Field | Rule |
|---|---|
image_url | Public HTTPS only; null clears it |
prompt | Required, 1 to 5000 characters |
negative_prompt | Omit or ""; non-empty returns 400 |
duration, duration_seconds | Rejected; steer length in the prompt |
mode | async, sync, subscribe, webhook |
Where to get the still
For a scene-matched score, the Music 1.0 docs suggest passing the accepted scene still as image_url when it fits, and varying genre family, tempo (at least 12 BPM apart) and lead instrument between contrasting scenes. If you want one consistent score, keep those fixed.
What not to expect
The docs describe image input as conditioning, not a guarantee that the audio will map to what is in the picture. Listen to the result and rewrite the brief, not the image, first.
Sources
Related posts
More in Models
- Inworld TTS 2 style steering and the Sume voiceover path
What is reported about Inworld TTS 2 style steering and 100+ languages, and how to produce a voiceover with Sume's tts_create tool and join takes.
- Nano Banana 3: does it exist? Current ids
Google autocomplete suggests a Nano Banana 3, but suggestions are not releases. What autocomplete lists and which Nano Banana id Sume accepts.
- Lip sync with a hand over the mouth: sync-3 vs Fabric on Sume
Sync Labs says sync-3 handles obstructions on faces. Sume's lip-sync routes start from a still plus audio. What each takes as input and what is promised.
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
Written by Sume