AI music for a tribute slideshow: a gentle bed for a memorial video

Score a tribute or memorial slideshow with AI music: an instrumental prompt, stills as timed holds, fades, and a listen-through before the video is shared.

5 min readSume
All posts

To score a tribute slideshow with AI music, generate a gentle instrumental bed with the Music Router, then lay your photos over it in a Timeline render. Sume treats a still as a static hold with a duration you choose, so each photo is one video[] slot, and the music goes in the soundtrack block with a fade-out. Ask for "Instrumental, no vocals" so no sung words appear that you did not write.

A memorial video is a case where a miss matters more than usual, so the steps here lean on listening. Sume cannot judge whether a track is right for a person, and no tool can. What it can do is give you several quiet takes cheaply, hold the photos for as long as you decide, and fade the music out softly.

What should the prompt ask for?

The Music docs ask for a brief with an emotion described precisely, a genre, tempo as a number, key and mode, two to four instruments with texture, and an arc with one named moment. They call these creative directions, not guaranteed settings. For a tribute that means choosing warmth and restraint on purpose. Here are example choices, which are yours to change.

A brief for a tribute bed, using the axes in the Music 1.0 docs, read 2026-10-03
AxisExample choiceReason
EmotionTender, grateful, quietly hopeful"Sad" alone tends toward generic; precise words steer better
GenreChamber folk or solo piano pieceFew parts, room for the photos
TempoA slow number such as 60 BPMGives each photo time on screen
InstrumentsPiano, soft strings, a single cello lineMid-range and warm
ArcOne gentle lift for the final photosA named moment, nothing abrupt
Closing clauseInstrumental, no vocalsKeeps the audio free of words you did not choose

How long should the track be?

If someone will speak over the slideshow, add "no spoken word" to the closing clause, as the docs advise for narration, so the track leaves space. Choose a length that fits the whole slideshow. Google's Lyria page puts a Lyria 3.5 track at up to three minutes, which is about thirty photos at six seconds each.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tribute-bed-001" \
  -d '{
    "prompt": "Tender, grateful and quietly hopeful solo piano with soft strings, 60 BPM, A-flat major. A single cello line enters at 1:00. One gentle lift for the last minute, then a slow, open ending. A 3-minute track. Instrumental, no vocals, no spoken word."
  }'

How do I build the slideshow?

Import your photos first. Timeline needs every source URL to be this workspace's media.sume.com artifact or asset, and off-host URLs are rejected, so use POST /v1/media-imports. Each slot needs a start, with the first at 0, and a duration of at least 0.2 seconds. Starts must increase, and the audio length is declared in audio.duration_seconds. Use audio.mode: "silence" with the bed in soundtrack, and set gain_db: 0, because Sume's compiler otherwise applies minus 16 dB to a bed that plays alone.

A fade transition of up to one second between photos is gentler than a hard cut, and fit: "blur" fills the frame around a photo that does not match the output shape. The default output is 1080x1920, so set width and height for a landscape video.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tribute-render-001" \
  -d '{
    "audio": {"mode": "silence", "duration_seconds": 24},
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/p1.jpg", "start": 0, "duration": 6, "fit": "blur"},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/p2.jpg", "start": 6, "duration": 6, "fit": "blur", "transition": {"type": "fade", "duration": 1}},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/p3.jpg", "start": 12, "duration": 6, "fit": "blur", "transition": {"type": "fade", "duration": 1}},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/p4.jpg", "start": 18, "duration": 6, "fit": "blur", "transition": {"type": "fade", "duration": 1}}
    ],
    "output": {"width": 1920, "height": 1080, "fade_in_seconds": 2, "fade_out_seconds": 3},
    "soundtrack": {"url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": 0, "fade_out_seconds": 6}
  }'

What should I listen for?

Before the video goes to anyone, play the finished file through, with the sound up, from start to end. Note anywhere the music turns brighter than the photos, anywhere a photo changes on a harsh beat, and the last ten seconds, which a family will remember. If a take is off, regenerate with a single change, and set a new Idempotency-Key, since the same key and payload return the original job.

A track that stops before the slideshow ends leaves silence, and the render adds a soundtrack_shorter_than_spine warning. Add loop: true if you would rather it repeat, or cut the slideshow to the track. Plan for four or five takes, at the fixed Music price each, and pick the one that sits quietly under your own photos.

  • The first ten seconds: does it arrive gently, or announce itself?
  • Photo changes: do they land on a harsh beat?
  • The ending: does it open out and settle, or stop?
  • Any voice-like synth that could be mistaken for words.
  • The level: bed alone at gain_db 0, or lower if the room is quiet.

What does a generated bed not do?

Sume cannot match the track to the person, score it to specific moments in a recording, or take a voice you provide and sing it. If a family member wants a particular song played, that is a licensing question for the song, and a generated bed is a stand-in, not a replacement.

A note on AI disclosure: Google's Lyria page says tracks carry a SynthID watermark. If the video will be shared somewhere that asks how audio was made, a line in the description is enough.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume