Christmas ad music: an instrumental Lyria brief that sells
A Christmas ad needs music that is festive without vocals under the voiceover. Brief it in one prompt, set the bed under speech and fade it out, with Sume.

What music suits a Christmas ad? An instrumental bed with a clear tempo, a recognisable seasonal color from instruments rather than a famous melody, and room for a voiceover. With Sume you describe that in one prompt to the Music Router, add "Instrumental, no vocals.", and let Timeline 1.0 handle the level and the ending.
The docs set the ground rules: the prompt is 1 to 5000 characters, length is steered in the prompt, duration is rejected, and a non-empty negative_prompt returns a 400, so every exclusion goes into the positive wording.
Write the brief for a voiceover
A bed under speech needs space in the middle of the sound. Ask for it by naming what plays and what stays out of the way. The seven-axis structure in the Music 1.0 docs gives you slots for each decision.
| Axis | Example for a Christmas ad |
|---|---|
| Emotion | Warm and hopeful, a little playful |
| Genre | Chamber pop with a jazz lean |
| Tempo | 96 BPM |
| Key and mode | G major with a lydian lift |
| Instruments | Celesta, upright bass, brushed drums, sleigh bells low in the mix |
| Arc | Sparse intro, the celesta melody enters at 0:06, full return at 0:22 |
| Era | 2024 hyper-clean |
Send the prompt
A synchronous request waits up to 30 seconds for a finished job; otherwise poll the job envelope.
import os
import requests
r = requests.post(
"https://api.sume.com/v1/music-router/generate",
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "christmas-ad-001",
},
json={
"model": "sume/music-auto",
"prompt": "Warm and hopeful chamber pop with a jazz lean, 96 BPM, G major with a lydian lift. Celesta, upright bass, brushed drums, sleigh bells low in the mix. Sparse intro, the celesta melody enters at 0:06, full return at 0:22. A 30-second track. Instrumental, no vocals.",
"mode": "sync",
"wait_timeout_seconds": 30,
},
timeout=60,
)
print(r.status_code)
print(r.json())Lay it under the voice
On the render request, the audio spine carries the voiceover and soundtrack carries the bed. duck_db lowers the bed under speech and takes values from 0 to 20, fade_out_seconds ends it softly (up to 10), and output.fade_in_seconds and fade_out_seconds handle the picture edges up to 5 seconds each. If a fade is longer than the audio, the render refuses it with soundtrack_fade_exceeds_output.
Check before you ship
Listen to the result and compare it to the brief. The docs state the axes are directions, not guaranteed settings, so plan a second pass. If the bed is almost right, change one axis, not all seven.
Sources
Related posts
More in Use cases
- Clean a transcript with a 2,000-character instruction, then caption
ElevenLabs speech to text can edit a transcript from a natural-language instruction up to 2,000 characters. A cleanup then captions workflow with Sume captions.
- Article 50: creator duties vs provider duties, side by side
Article 50(4) puts deepfake and certain text disclosure on deployers, while 50(2) puts marking on providers. A two-column split for teams that publish AI video.
- Cyber Week five-day creative sprint: a daily Seedance clip budget
Shopify defines BFCM as Thanksgiving through Cyber Monday. Twenty-two 8-second 720p vertical clips across those five days cost $101.6928 on Sume.
- Demand Gen carousels: 2 to 10 matching cards from one reference
Demand Gen carousels take 2 to 10 cards. Image assets run 4:5 or 9:16 at 5 MB. How to batch matching cards from one reference image with the Sume image API.
Written by Sume