Music for an avatar video: pass the preview still as image_url

Approve an avatar preview, then send its preview_image_url to the Music Router as image_url so the track matches the frame. Steps, limits and the $0.125 cost.

5 min readSume
All posts

To score an avatar video from its own look, create an avatar video preview, read its preview_image_url, and pass that still as image_url to POST /v1/music-router/generate along with a written brief. The music request is one fixed-price generation of $0.125 (read 2026-10-03), and the preview step lets you settle framing before you pay for the full render.

Image conditioning is optional on Sume Music, and the docs describe it as visual conditioning on top of a text prompt. It does not replace the brief: you still have to say tempo, mood, instruments and 'instrumental, no vocals'.

The flow in four steps

The preview docs say the stills are public-safe fields, and the Music docs say image_url must be a public HTTPS image. Check that the URL you pass is fetchable from outside your workspace; if the Music request rejects it, fall back to describing the frame in words.

  • Create the preview with POST /v1/avatar-video-previews using your script, scene and aspect_ratio.
  • Poll the job, then read the preview resource: preview_image_url is the primary still and scene_previews[] has one per scene.
  • Send the still to the Music Router as image_url, with a seven-axis brief written for the mood of the frame.
  • Approve the preview, call generate-video, and combine the music with the finished clip in Timeline.

What the request looks like

Notice that the prompt gives the structure and the image gives the mood. Do not send duration or duration_seconds, which are rejected, and do not send a non-empty negative_prompt. Ask for the length in the prompt, for example 'a 20-second track'.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-bed-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Warm minimal electronic underscore, 96 BPM, A minor. Soft pads, a muted pluck motif, light brushed percussion entering at 0:08. A 20-second track. Instrumental, no vocals.",
    "image_url": "https://media.sume.com/artifacts/your-preview-still.png"
  }'

Cost and trade-offs

Music is a fixed price whether or not you attach an image, so conditioning on the still adds nothing to the bill. The trade-off is time: you now have a dependency between two jobs. If you do not need the match to be tight, generate music in parallel from text alone and pick between candidates.

Avoid music that fights narration. A presenter's speech occupies the mid frequencies, so ask for sparse mids and keep the bed quiet in the Timeline mix. The docs advise adding 'no spoken word' only when music sits under narration, so that the track does not contain its own voice.

Steps and costs, read 2026-10-03
StepWhat you getCost note
Avatar previewFirst-frame stillsSee the avatar video preview docs
Music generation$0.125 fixedSame with or without image_url
Final avatar renderPer second by tier$0.245 on Plus, no product
Timeline join$0.10 per output minuteNeeds Sume-hosted media

A cost check for the whole chain

The chain has four paid parts. The music generation is $0.125 whatever its inputs. The avatar video is per second: a 20-second Plus clip is $4.90 and Standard $3.68. The Timeline join is $0.10 per output minute, so a 20-second video rounds up to one minute at most. Together that is about $5.125 for a finished 20-second Plus video with a matched score, before the avatar is created and any retakes.

The preview stage is the only step whose price I did not find on the rate card, so read your usage after the first preview and add that line to your own sheet.

When the image is the wrong input

If the avatar preview shows a plain studio, the still tells the model very little about mood, and the text prompt does all the work. In that case skip the image and spend the effort on the brief. Image conditioning pays off when the scene has a distinct look, such as a warm kitchen, a neon street or a bright product shot, that you want the music to echo.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume