Explainer music that changes per scene: three Lyria briefs, one concat

Make each scene's music a separate Lyria 3.5 brief at least 12 BPM apart, then join them gaplessly with Timeline audio concat for $0.01.

5 min readSume
All posts

For an explainer where each scene needs different music, make one Lyria 3.5 track per scene with POST /v1/music-router/generate, then join them with operation: "concat" on POST /v1/timeline-1.0/audio. The concat is gapless, sample-domain and costs $0.01 flat. Steer each track's length in its prompt, because duration and duration_seconds are rejected.

Source: Music Router and Timeline audio, read 2026-10-04.

Three briefs

Sume's music brief template says contrasting scenes should vary genre family, tempo by at least 12 BPM, and lead instrument. Three briefs below follow that.

Scene briefs for a 90-second explainer (read 2026-10-04)
ScenePrompt summaryBPMLead
ProblemTense minimal electronic, a 30-second track78Pulsing synth bass
SolutionBright acoustic pop, a 30-second track102Plucked guitar
Call to actionConfident funk, a 30-second track118Brass stabs

Submit the first brief

The BPM gaps are 24 and 16, both above the 12 BPM rule. Put exclusions in the positive prompt, since a non-empty negative_prompt is rejected. Each request is one job, so use an Idempotency-Key per scene.

curl -X POST https://api.sume.com/v1/music-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: explainer-music-problem-001" \
  -d '{
    "model": "sume/music-auto",
    "prompt": "Tense minimal electronic, 78 BPM, A minor. Pulsing synth bass, sparse clicks, no drums until 0:20. A 30-second track. Instrumental, no vocals."
  }\

Join them

When all three tracks are ready, read each audio artifact from result.artifacts[]. Tracks must be on media.sume.com and share a channel layout, or concat fails with audio_parts_channel_mismatch. Concat takes 1 to 20 parts.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: explainer-music-concat-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/problem.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/solution.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/cta.wav" }
    ]
  }\

Line up the picture

Lyria does not guarantee exactly 30 seconds, so read segments[] from the concat result. It gives each part's start and duration_seconds, which are the offsets to re-base Timeline 1.0 video[].start against. Use the merged file as audio.url on the render.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume