Score a scene from its still: image_url on Sume music, Lyria 3.5

Pass the accepted scene still as image_url on a Sume music request so the score matches the picture. What the field accepts, what it does not, and the cost.

4 min readSume
All posts

To score a scene from its picture, pass the approved still as image_url on the Sume music request and describe the music in the prompt. The field takes one public HTTPS image, and the image conditions the result alongside your text; it does not replace the brief. Sume's Music docs say to pass the accepted scene still when you want continuity between picture and score.

Google's Gemini API changelog (read 2026-10-02) lists text and image inputs for Lyria 3.5, which went GA on 2026-09-03. Sume's endpoint exposes a single image URL, so this post covers what you can do with one.

What does image_url accept?

One public HTTPS URL. Links that are not public HTTPS are not accepted. You can also send null to clear an image on a client that reuses request objects. Use a still you already approved, ideally a Sume media.sume.com artifact, so the URL stays reachable.

music request fields that matter for a scene score. Sume docs, read 2026-10-02.
FieldRequiredNote
promptYes1 to 5000 characters; exclusions go here
image_urlNoPublic HTTPS image, or null to clear
negative_promptNoUnsupported when non-empty
modeNoasync, sync, subscribe or webhook

What should the prompt still say?

Say the music, not the picture. The still gives mood; the prompt gives tempo, key, instruments and the arc. A prompt of the form cinematic ambient underscore matching the mood of the reference still, instrumental only is the docs' own example, and you should add a number for tempo to it.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: scene-score-001" \
  -d '{"prompt": "Warm cinematic underscore, 80 BPM, F major. Felt piano, soft strings. Instrumental only.", "image_url": "https://example.com/scene-03.png"}'

What if the score does not match the picture?

Treat the output as a take. There is no seed, so change the brief, not a setting, and try again. Verify the audio by listening; the result's lyrics field is model-reported metadata about tempo and structure, not an audio measurement.

What does it cost?

A fixed $0.125 per accepted generation on the Music 1.0 page; the price does not vary with prompt length or image conditioning.

Which stills work well?

Pick a still that shows mood: light, color, setting and pace. A frame full of small text tells the model little about music. If your scenes share a look, pass the same still family across them so the scores feel related, and change the prompt's tempo and instruments for each scene.

Make sure the URL stays reachable: use a stable public address rather than a link that may expire before the job runs.

Should I use a still on every scene?

Not necessarily. Use one when the picture carries a mood you want the music to share, and skip it when the prompt alone says what you need. The image is optional, and passing null clears it on a client that reuses request objects.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume