Score a scene from its still: image_url on Sume music, Lyria 3.5
Pass the accepted scene still as image_url on a Sume music request so the score matches the picture. What the field accepts, what it does not, and the cost.

To score a scene from its picture, pass the approved still as image_url on the Sume music request and describe the music in the prompt. The field takes one public HTTPS image, and the image conditions the result alongside your text; it does not replace the brief. Sume's Music docs say to pass the accepted scene still when you want continuity between picture and score.
Google's Gemini API changelog (read 2026-10-02) lists text and image inputs for Lyria 3.5, which went GA on 2026-09-03. Sume's endpoint exposes a single image URL, so this post covers what you can do with one.
What does image_url accept?
One public HTTPS URL. Links that are not public HTTPS are not accepted. You can also send null to clear an image on a client that reuses request objects. Use a still you already approved, ideally a Sume media.sume.com artifact, so the URL stays reachable.
| Field | Required | Note |
|---|---|---|
| prompt | Yes | 1 to 5000 characters; exclusions go here |
| image_url | No | Public HTTPS image, or null to clear |
| negative_prompt | No | Unsupported when non-empty |
| mode | No | async, sync, subscribe or webhook |
What should the prompt still say?
Say the music, not the picture. The still gives mood; the prompt gives tempo, key, instruments and the arc. A prompt of the form cinematic ambient underscore matching the mood of the reference still, instrumental only is the docs' own example, and you should add a number for tempo to it.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scene-score-001" \
-d '{"prompt": "Warm cinematic underscore, 80 BPM, F major. Felt piano, soft strings. Instrumental only.", "image_url": "https://example.com/scene-03.png"}'What if the score does not match the picture?
Treat the output as a take. There is no seed, so change the brief, not a setting, and try again. Verify the audio by listening; the result's lyrics field is model-reported metadata about tempo and structure, not an audio measurement.
What does it cost?
A fixed $0.125 per accepted generation on the Music 1.0 page; the price does not vary with prompt length or image conditioning.
Which stills work well?
Pick a still that shows mood: light, color, setting and pace. A frame full of small text tells the model little about music. If your scenes share a look, pass the same still family across them so the scores feel related, and change the prompt's tempo and instruments for each scene.
Make sure the URL stays reachable: use a stable public address rather than a link that may expire before the job runs.
Should I use a still on every scene?
Not necessarily. Use one when the picture carries a mood you want the music to share, and skip it when the prompt alone says what you need. The image is optional, and passing null clears it on a client that reuses request objects.
Sources
Related posts
More in Use cases
- Translate text inside an image: one Nano Banana 2 edit per language
Google says Nano Banana 2 can translate and localize text within an image. Loop one edit per language through Sume POST /v1/images and keep the layout fixed.
- Nano Banana Pro 4:5 is not 1080x1350: crop and resize in Python
On Sume, Nano Banana Pro takes 4:5 at about 928x1152, not exact 1080x1350. Resize to Instagram's size locally with Pillow, with the small crop explained.
- Narrate a 2,000-word blog post with an AI voice: cost and steps
A 2,000-word post is roughly 12,000 characters, so one Sume TTS job at $0.0475 per 1,000 characters. The steps, settings and what to check before publishing.
- New York FAIR News Act: AI video labels and what is pending
New York's FAIR News Act would require a label on news video substantially made by generative AI. It awaits the governor; here is what the sources say.
Written by Sume