Corporate onboarding video music: a 5-minute ducked bed
Add a 5-minute instrumental bed under onboarding narration with duck_db in Timeline: one generation and a 5-minute render, $0.625 on Sume.

For a 5-minute onboarding video, generate one neutral, upbeat instrumental of about five minutes and mix it under the narration with duck_db in Timeline. On Sume that is one generation at $0.125 and a five-minute render at $0.50, so $0.625 for the audio and assembly.
Onboarding content is long and informational. A bed that is too lively tires the viewer, so the brief aims at quiet optimism with very little movement and no vocals. A long piece also benefits from named sections so the music has some shape without a distracting climax.
The brief to send
The sections below give a soft intro, a steady middle, and a light close, each with its own small change.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: corporate-onboarding-video-music-5-minut-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Friendly corporate onboarding bed, 100 BPM, C major. Soft electric piano, light acoustic guitar, gentle bass, minimal soft percussion. [0:00-0:30] Intro. [0:30-4:30] Steady groove with slow variation. [4:30-5:00] Close. Optimistic and neutral, clean modern production. Instrumental, no vocals. A 5-minute track."
}'Why each axis is set that way
| Axis | Choice |
|---|---|
| Emotion | Optimistic, neutral, supportive |
| Genre | Light corporate pop-acoustic |
| Tempo | 100 BPM, steady |
| Key and mode | C major, plain |
| Instruments | Electric piano, acoustic guitar, bass, soft percussion |
| Arc | Soft intro, steady middle, light close |
| Era and production | Clean, modern, mid-bright |
Getting the length right
Ask for five minutes in the prompt and verify the real length. With the narration as the spine, set soundtrack.duck_db so the bed dips while someone speaks; ducking is rejected on a silent spine. Use soundtrack.gain_db for the overall bed level and fade_out_seconds for the close.
If the narration runs longer than the track, soundtrack.loop repeats it. A five-minute chapter is easier to produce than a 20-minute video, so split a long course into chapters and render each separately.
Before paying for a render, call POST /v1/timeline-1.0/plan with the same body. It is unbilled and reports the billable minutes, which for this 300-second video is 5. Then send the real POST /v1/timeline-1.0/render with an Idempotency-Key, wait up to 30 seconds in sync mode, and poll GET /v1/jobs/:id/status if it is still queued or processing. Statuses are queued, processing, completed, failed and canceled, and a 429 queue_full means try again shortly.
What it costs
A five-minute render is five started minutes, $0.50. Splitting the video into five one-minute modules would cost the same $0.50 in render, but five generations, $0.625 for music instead of $0.125, so one bed per video is cheaper than one bed per module. Scaling is simple arithmetic. One finished video here is $0.625 of audio and assembly, so ten of them are $6.25 and fifty are $31.25, before any retries. Failed jobs are not billed as accepted generations, but a regeneration you ask for is a new job, so budget one or two extra takes for a hero clip.
| Line | Rate | Amount |
|---|---|---|
| Music generation x 1 | $0.125 each | $0.125 |
| Timeline render, 300 s = 5 started minutes | $0.10 per started minute | $0.50 |
| Total for the audio and assembly | $0.625 |
Check before you publish
Use one brand bed across the onboarding series. Because generation is flat-priced, you could also give each department its own variant for the cost of $0.125 each.
Keep the prompt in your style guide so new videos sound consistent.
The hand-off between the two calls is the step people miss. When the music job completes, GET /v1/jobs/:id/result returns result.artifacts[] with an audio/mpeg file on media.sume.com. Timeline only accepts workspace media URLs, so import any outside file through POST /v1/media-imports first, and send the resulting URL as soundtrack.url. The Music docs also note that job.request.routed_model names the engine that actually ran, which is worth saving with the file.
- Do not use music to carry policy information; the narration does that.
- Use captions for compliance content and keep the bed low.
- Check the final level on laptop speakers, which is how most people watch.
- Confirm usage terms for generated music with Sume before internal or external distribution.
Sources
Related posts
More in Media tools
- Crop 21:9 to 9:16: a 608 px strip, then scale to 1080x1920
Cropping 2560x1080 ultrawide to 9:16 leaves a 608 px wide strip. The Sume video filter crop fractions, the scale step to 1080x1920, and the sharpness trade-off.
- Crop 2560x1080 ultrawide to 16:9: video filter width 0.75
A 21:9 clip at 2560x1080 crops to exactly 1920x1080 with one Sume video filter crop op, width 0.75 and x 0.125. The math, the request, and rounding.
- Crop a 16:9 AI clip to 21:9 with video filter: height 0.762, y 0.119
Cut a 16:9 Wan 3.0 or Kling 3 clip to a 21:9 frame for $0.02: crop width 1, height 0.762, y 0.119. Compared with generating 21:9 natively on Sume.
- Cut a 90 s music window: timeline-audio split and range end errors
Use timeline-audio split with ranges [{start, end}] to cut a window. end must exceed start or you get audio_range_end_before_start; omit end for the rest.
Written by Sume