Score a six-image moodboard: six music beds from stills for $0.75

Music 1.0 and the Music Router take an optional image_url. Six stills, six beds, $0.125 each: $0.75. What the image does, what it does not, and the limits.

5 min readSume
All posts

To score a six-image moodboard, send six Music Router requests, each with the still as image_url and a short brief in prompt. Each accepted generation costs a fixed $0.125, so six beds cost $0.75. The image conditions the music; it does not set its length, tempo or key, and the Music docs call the brief's axes creative directions, not guaranteed output values.

What image conditioning means here

Sume's Music docs describe the request as text-first with optional image conditioning. The only image field is image_url, which must be public HTTPS; send null only if a client that reuses request objects needs to clear it. The price does not change with the image or the prompt length.

The docs also say to pass the accepted scene still as image_url when scenes in one project contrast and you want each bed to match its picture, or when you want one consistent score and keep continuity. For a moodboard that is the natural use: one still, one bed.

Cost of scoring a moodboard, rates read 2026-10-09
StillsMusic jobsPrice eachTotal
11$0.125$0.125
66$0.125$0.75
1212$0.125$1.50
6, with one retry on 2 stills8$0.125$1.00

Writing the prompt around a picture

Do not describe the image in the prompt; the model already sees it. Spend the words on what a picture cannot say: tempo as a number, key and mode, two to four instruments with texture, and one named moment. The docs list seven axes (emotion, genre lineage, tempo, key and mode, instruments, arc, era) and advise ending with one clause: "Instrumental, no vocals."

Length is also prompt text. Music 1.0 has no duration field and rejects duration and duration_seconds, so write "a 20-second track" and measure the file you get back.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: moodboard-still-03" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Hushed, slightly melancholic, 72 BPM, D minor. Rhodes through tape wow, brushed snare, one muted trumpet answer at 0:12. A 20-second track. Instrumental, no vocals.",
    "image_url": "https://example.com/moodboard/03.png"
  }'

From six beds to a finished cut

A bed is rarely the exact length of the shot it sits under. In a Timeline render the optional soundtrack object takes the bed's url, a gain_db, a loop flag, a fade_out_seconds up to 10 and a duck_db from 0 to 20 that lowers the bed under speech; ducking needs a real audio spine, not silence. A 20-second bed under a 45-second voiced clip can loop with a fade, which is cheaper than a second generation at $0.125.

If a bed is too long instead, a Timeline audio split range cuts it for $0.01 per job. Beds you keep for later are plain media.sume.com audio files, so they can be reused across renders without paying again.

Limits worth knowing before you batch

Four details decide whether a six-job batch behaves.

  • The prompt is 1 to 5,000 characters and exclusions go in the positive text. A non-empty negative_prompt returns 400 with public_reason=negative_prompt_unsupported.
  • Use a distinct Idempotency-Key per still. Reusing one key returns the first job, which is useful for retries and wrong for a batch.
  • Read the audio from result.artifacts[] where type is audio. Raw provider URLs are not public outputs.
  • If you want to know which engine ran, job.request.routed_model names it, for example lyria-3.5. The Music Router charges the same fixed price on every catalog model.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume