AI classroom background music for lesson videos, under narration
Make calm instrumental music for a lesson video: a prompt that keeps vocals out, a Python script, and a Timeline bed that ducks under the teacher's voice.

To make classroom background music with AI, ask the Music Router for an instrumental bed in the prompt, then lay it under the lesson narration with a Timeline render that ducks the music while the teacher speaks. On Sume that is two jobs: one POST /v1/music-router/generate for the track and one POST /v1/timeline-1.0/render for the mix. Length is steered in the prompt, because the music API rejects duration fields.
Lesson videos have one demand most music does not: nothing in the track may compete with spoken words. That rules out vocals, busy drum fills and bright lead lines, and it makes the bed's closing clause matter. The rest of this post covers the prompt, a script that fetches the track, and the mix.
What should the prompt say for a classroom bed?
Sume's Music docs say to close the brief with one clause, "Instrumental, no vocals", and to add "no spoken word" when there is narration. Both are exclusions written into the positive prompt, because a non-empty negative_prompt returns HTTP 400. The brief template uses seven axes: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era. The docs call them creative directions, not guaranteed settings, so listen to the result.
| Axis | Lesson-bed choice (example) | Why it suits narration |
|---|---|---|
| Emotion | Calm, curious, unhurried | Does not pull attention off the speaker |
| Tempo | A slow number such as 72 BPM | A stated number is a stronger hint than "slow" |
| Instruments | Soft piano, light acoustic guitar, quiet pad | Few parts and mid-range textures |
| Arc | One gentle lift at the midpoint, nothing sharp | Avoids a moment that fights a key sentence |
| Closing clause | Instrumental, no vocals, no spoken word | Stops sung or spoken parts competing with the lesson |
How do I fetch the track with Python?
This script submits the job, polls the status URL until the job is terminal, and prints the audio artifact URL. It reads terminal, sume_status and next_poll_after_seconds from the status response, as the Jobs and results page describes.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def data(r):
r.raise_for_status()
body = r.json()
return body.get("data", body)
prompt = ("Calm, curious instrumental bed for a classroom lesson, 72 BPM, "
"C major. Soft piano, light acoustic guitar, quiet pad. One gentle "
"lift at the midpoint. A 3-minute track. Instrumental, no vocals, "
"no spoken word.")
job = data(requests.post(f"{API}/v1/music-router/generate",
headers={**H, "Idempotency-Key": "lesson-bed-001"},
json={"prompt": prompt}))
job_id = job.get("job_id") or job.get("request_id") or job["id"]
while True:
s = data(requests.get(f"{API}/v1/jobs/{job_id}/status", headers=H))
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 5)
if s["sume_status"] != "completed":
raise SystemExit(f"job ended {s['sume_status']}")
res = data(requests.get(f"{API}/v1/jobs/{job_id}/result", headers=H))
print([a["url"] for a in res["result"]["artifacts"] if a["type"] == "audio"][0])How do I mix it under the teacher's voice?
Timeline 1.0 takes one audio spine, which here is the narration, plus ordered video[] slots and an optional soundtrack. The soundtrack block takes the music url, gain_db, loop, fade_out_seconds up to 10, and duck_db from 0 to 20. Ducking needs a real spine, so it works with narration and is refused over silence (duck_requires_audio_spine). Every URL must already be a media.sume.com artifact or asset, so import a narration file you recorded elsewhere first with POST /v1/media-imports.
A long lesson outlasts a music track, since Google's Lyria page puts a Lyria 3.5 track at up to three minutes. Set loop: true so the bed repeats, and the renderer will not warn about a short bed. A render can run up to 1800 seconds. The values below are starting points to tune by ear, not recommendations from the docs.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lesson-mix-001" \
-d '{
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/narration.wav", "duration_seconds": 300},
"video": [{"source_url": "https://media.sume.com/artifacts/artf_demo/slides.mp4", "start": 0, "duration": 300}],
"soundtrack": {"url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
"gain_db": -14, "loop": true, "duck_db": 12, "fade_out_seconds": 6}
}'How do I choose between takes?
Play each take under a real recording of the lesson before you commit, not on its own. A bed that sounds pleasant alone often turns out to have a piano line that lands on top of consonants, or a pad that muddies the room tone. Three checks are enough: can you follow every sentence at normal volume, does the bed still feel present in the pauses, and does the opening of the loop match its ending closely enough to repeat.
If a take fails the first check, lower gain_db before you regenerate. If it keeps failing, rewrite the instruments line to fewer, softer parts, since the music docs ask for two to four instruments with texture. Keep the closing clause unchanged. Google's Lyria page says tracks carry a SynthID watermark, which is worth mentioning if your school or district asks how the audio was made.
What does a lesson bed cost, and what are the limits?
Costs are in the docs. A Music Router generation is a fixed price, $0.125 per accepted generation on the Music 1.0 page, and a render is $0.10 per output minute rounded up. A five-minute lesson render is therefore five of those minutes, plus the one music job.
Two limits are worth knowing before you build a course library. The loop is a plain repeat, so a bed with a strong ending can sound abrupt each time it restarts; ask for a soft, open ending in the brief. And the default output is a 1080x1920 portrait MP4, so set output.width and output.height to even numbers such as 1920 and 1080 for a landscape lesson.
Sources
Related posts
More in Use cases
- AI person in a Meta ad: is the label next to Sponsored?
Meta puts AI info next to Sponsored when its own tools make a photorealistic human. For outside tools it describes About this ad. How to add your own cue.
- AI image slideshow Shorts on YouTube: what narrative to add
YouTube's inauthentic content policy lists image slideshows with minimal narrative as not allowed. Turn stills into a Short with a real story, using Sume.
- Ad music cutdowns: 15, 30 and 60 seconds from one AI track
Make 15, 30 and 60 second versions of one AI music bed: three prompts or one cut with a fade-out. Costs and limits from Sume's Music Router and Timeline docs.
- AI music for a tribute slideshow: a gentle bed for a memorial video
Score a tribute or memorial slideshow with AI music: an instrumental prompt, stills as timed holds, fades, and a listen-through before the video is shared.
Written by Sume