Graduation video music: a 2-minute anthem with a build
Brief a 2-minute graduation video bed with a build to a triumphant moment at 1:20 for the diploma reel: one generation and a render, $0.325 on Sume.

For a graduation video, ask for a 2-minute uplifting orchestral-pop anthem that builds to a triumphant lift at 1:20, where the diploma reel or the class photo lands. On Sume that is one generation at $0.125 and a two-minute render at $0.20, so $0.325.
A graduation video is emotional and has an obvious peak. The brief should therefore have a real arc: quiet start for the memories, rising middle, a clear climax and a warm resolution. Name the climax time, because that is the anchor your edit needs.
The brief to send
The sections below give a slow piano opening, a steady build with percussion, and the lift at 1:20.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: graduation-video-music-a-2-minute-anthem-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Uplifting graduation anthem, 92 BPM, E-flat major. Piano, warm strings, snare build, French horns, light cymbal swells. [0:00-0:40] Gentle piano memories. [0:40-1:20] Steady build. [1:20-1:50] Triumphant lift. [1:50-2:00] Warm resolve. Cinematic pop production. Instrumental, no vocals. A 2-minute track."
}'Why each axis is set that way
| Axis | Choice |
|---|---|
| Emotion | Proud, nostalgic, uplifting |
| Genre | Orchestral pop |
| Tempo | 92 BPM, with a snare build |
| Key and mode | E-flat major, bright |
| Instruments | Piano, strings, horns, snare, cymbals |
| Arc | Build from 0:40, lift at 1:20, resolve at 1:50 |
| Era and production | Cinematic pop, polished |
Getting the length right
Section markers are the way to steer structure in Lyria 3.5. The router rejects duration and duration_seconds, so the two minutes are in the prompt. Check where the lift actually lands in the returned file and align the diploma reel to it.
Set audio.duration_seconds to 120 in the render and fade_out_seconds to 2 or 3 for a smooth finish. If a speaker's voice is part of the video, make that the spine and set duck_db.
Before paying for a render, call POST /v1/timeline-1.0/plan with the same body. It is unbilled and reports the billable minutes, which for this 120-second video is 2. Then send the real POST /v1/timeline-1.0/render with an Idempotency-Key, wait up to 30 seconds in sync mode, and poll GET /v1/jobs/:id/status if it is still queued or processing. Statuses are queued, processing, completed, failed and canceled, and a 429 queue_full means try again shortly.
What it costs
A two-minute render is $0.20. Per-class versions, for example three schools with their own mood, are three generations at $0.375 plus three renders at $0.60, which is $0.975. The cost scales with the number of videos, not their length. Scaling is simple arithmetic. One finished video here is $0.325 of audio and assembly, so ten of them are $3.25 and fifty are $16.25, before any retries. Failed jobs are not billed as accepted generations, but a regeneration you ask for is a new job, so budget one or two extra takes for a hero clip.
| Line | Rate | Amount |
|---|---|---|
| Music generation x 1 | $0.125 each | $0.125 |
| Timeline render, 120 s = 2 started minutes | $0.10 per started minute | $0.20 |
| Total for the audio and assembly | $0.325 |
Check before you publish
Seasonal spikes happen in May and June in many places; prepare the briefs early.
Keep the best-performing brief for next year's class.
The hand-off between the two calls is the step people miss. When the music job completes, GET /v1/jobs/:id/result returns result.artifacts[] with an audio/mpeg file on media.sume.com. Timeline only accepts workspace media URLs, so import any outside file through POST /v1/media-imports first, and send the resulting URL as soundtrack.url. The Music docs also note that job.request.routed_model names the engine that actually ran, which is worth saving with the file.
- Do not name the school or students in the music prompt; put names on screen.
- Check that the lift lines up with the reel before final export.
- Provide a version without music if the venue plays its own.
- Confirm usage terms for generated music with Sume before distributing.
Sources
Related posts
More in Media tools
- Grayscale value check for AI product images: does it read at a glance?
Convert a Sume image to grayscale, measure the contrast between product and background with ImageStat, and flag low-value-contrast images before they ship.
- H3 Max Recast: 1-4 person photos, 768p default, 1080p opt-in
H3 Max Recast swaps the people in one 5-30 s video using 1-4 reference photos, one per new person. Output is 768p by default; 1080p is opt-in.
- Hailuo 3.0 clips in a Sume timeline: set output.fps or omit it
Cutting 24 fps AI clips into a timeline at the wrong frame rate causes judder. Sume's Timeline output.fps accepts 24, 25, 30 or 60, or matches your sources.
- Hand an AI clip to an editor: MP4, PNG frames and WAV on Sume
What Sume can deliver for a finishing editor: MP4 clips, lossless PNG stills at source size, and sample-exact WAV audio. No EXR or HDR is documented.
Written by Sume