Music for an avatar video: pass the preview still as image_url
Approve an avatar preview, then send its preview_image_url to the Music Router as image_url so the track matches the frame. Steps, limits and the $0.125 cost.
To score an avatar video from its own look, create an avatar video preview, read its preview_image_url, and pass that still as image_url to POST /v1/music-router/generate along with a written brief. The music request is one fixed-price generation of $0.125 (read 2026-10-03), and the preview step lets you settle framing before you pay for the full render.
Image conditioning is optional on Sume Music, and the docs describe it as visual conditioning on top of a text prompt. It does not replace the brief: you still have to say tempo, mood, instruments and 'instrumental, no vocals'.
The flow in four steps
The preview docs say the stills are public-safe fields, and the Music docs say image_url must be a public HTTPS image. Check that the URL you pass is fetchable from outside your workspace; if the Music request rejects it, fall back to describing the frame in words.
- Create the preview with
POST /v1/avatar-video-previewsusing your script, scene andaspect_ratio. - Poll the job, then read the preview resource:
preview_image_urlis the primary still andscene_previews[]has one per scene. - Send the still to the Music Router as
image_url, with a seven-axis brief written for the mood of the frame. - Approve the preview, call
generate-video, and combine the music with the finished clip in Timeline.
What the request looks like
Notice that the prompt gives the structure and the image gives the mood. Do not send duration or duration_seconds, which are rejected, and do not send a non-empty negative_prompt. Ask for the length in the prompt, for example 'a 20-second track'.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-bed-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Warm minimal electronic underscore, 96 BPM, A minor. Soft pads, a muted pluck motif, light brushed percussion entering at 0:08. A 20-second track. Instrumental, no vocals.",
"image_url": "https://media.sume.com/artifacts/your-preview-still.png"
}'Cost and trade-offs
Music is a fixed price whether or not you attach an image, so conditioning on the still adds nothing to the bill. The trade-off is time: you now have a dependency between two jobs. If you do not need the match to be tight, generate music in parallel from text alone and pick between candidates.
Avoid music that fights narration. A presenter's speech occupies the mid frequencies, so ask for sparse mids and keep the bed quiet in the Timeline mix. The docs advise adding 'no spoken word' only when music sits under narration, so that the track does not contain its own voice.
| Step | What you get | Cost note |
|---|---|---|
| Avatar preview | First-frame stills | See the avatar video preview docs |
| Music generation | $0.125 fixed | Same with or without image_url |
| Final avatar render | Per second by tier | $0.245 on Plus, no product |
| Timeline join | $0.10 per output minute | Needs Sume-hosted media |
A cost check for the whole chain
The chain has four paid parts. The music generation is $0.125 whatever its inputs. The avatar video is per second: a 20-second Plus clip is $4.90 and Standard $3.68. The Timeline join is $0.10 per output minute, so a 20-second video rounds up to one minute at most. Together that is about $5.125 for a finished 20-second Plus video with a matched score, before the avatar is created and any retakes.
The preview stage is the only step whose price I did not find on the rate card, so read your usage after the first preview and add that line to your own sheet.
When the image is the wrong input
If the avatar preview shows a plain studio, the still tells the model very little about mood, and the text prompt does all the work. In that case skip the image and spend the effort on the brief. Image conditioning pays off when the scene has a distinct look, such as a warm kitchen, a neon street or a bright product shot, that you want the music to echo.
Sources
Related posts
More in Use cases
- Nano Banana 2 catalog photos: draft at 0.5K, finish at 2K
Cut catalog image cost by drafting at 0.5K and rendering only approved shots at 2K with Nano Banana 2; vendor rates and the Sume resolution tiers.
- Narrate a 3-minute Short: Sume TTS, word timings and 60 s caption cuts
YouTube allows three-minute Shorts. A plan with Sume: one TTS request under 20,000 characters, word timings, and caption jobs in pieces of 60 seconds or less.
- Narrate an 80,000-word novel with AI TTS: $23 in 34 requests
An 80,000-word novel is about 480,000 characters, quoted at $23.00 on Sume TTS in 34 requests of up to 14,400 characters, roughly 10 hours of audio.
- Neighborhood guide video for real estate agents: photos and music
A 30-second neighborhood guide built from your own photos of local spots: stills in Timeline 1.0, one caption cue per spot, and a soundtrack bed.
Written by Sume