Yoga class video music: loop a 2-minute bed for 20 minutes
Generate one calm 2-minute ambient bed and loop it under a 20-minute yoga video with Timeline: one generation plus a 20-minute render, $2.125 on Sume.

For a 20-minute yoga class video, generate one calm 2-minute ambient bed and let Timeline loop it under the session with soundtrack.loop. On Sume that is one generation at $0.125 and a 20-minute render at $0.10 per started minute ($2.00), so $2.125 for the soundtrack and assembly.
Yoga music has a different job from ad music: no beats, no arc, no surprises. The point is steadiness, so the brief asks for a slowly evolving pad with no rhythm and a seamless feel at both ends, which is what makes looping unnoticeable.
The brief to send
The brief removes every cue that would reveal a loop: no drums, no melody, no cadence. It asks for even texture start to finish, and a quiet ending that matches the opening.
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: yoga-class-video-music-loop-a-2-minute-a-001" \
-d '{
"model": "sume/music-auto",
"prompt": "Calm yoga ambient, no rhythm, D major drone. Warm analog pad, soft singing bowl at low volume, distant flute breaths, slow filter movement. Even and steady from start to finish, no climax, start and end on the same soft texture. Spacious natural reverb. Instrumental, no vocals. A 2-minute track."
}'Why each axis is set that way
| Axis | Choice |
|---|---|
| Emotion | Peaceful, grounded, even |
| Genre | Ambient drone, meditation |
| Tempo | No pulse, free time |
| Key and mode | D major, a single sustained centre |
| Instruments | Analog pad, singing bowl, distant flute |
| Arc | None by design; same texture at both ends |
| Era and production | Spacious, natural reverb |
Getting the length right
Music Router has no length field, so ask for two minutes in the prompt and expect an approximate result. In Timeline, set audio.duration_seconds to 1200 and pass the MP3 as soundtrack.url with loop set to true. The loop repeats the track to cover the spine, and fade_out_seconds closes the session gently.
Timeline accepts audio.duration_seconds from 1 to 1800, which is 30 minutes, so a 20-minute class fits in one render. For a longer practice, render two parts and join them with the audio concat endpoint.
Before paying for a render, call POST /v1/timeline-1.0/plan with the same body. It is unbilled and reports the billable minutes, which for this 1200-second video is 20. Then send the real POST /v1/timeline-1.0/render with an Idempotency-Key, wait up to 30 seconds in sync mode, and poll GET /v1/jobs/:id/status if it is still queued or processing. Statuses are queued, processing, completed, failed and canceled, and a 429 queue_full means try again shortly.
What it costs
Here length is the cost driver, not music. The generation is fixed at $0.125 whatever the track length, while a render is billed per started output minute, so a 20-minute video is 20 minutes of render. A 30-minute class would be $3.00 plus the music. Looping one short track is the cheap way to cover a long video, because you pay for one generation, not ten. Scaling is simple arithmetic. One finished video here is $2.125 of audio and assembly, so ten of them are $21.25 and fifty are $106.25, before any retries. Failed jobs are not billed as accepted generations, but a regeneration you ask for is a new job, so budget one or two extra takes for a hero clip.
| Line | Rate | Amount |
|---|---|---|
| Music generation x 1 | $0.125 each | $0.125 |
| Timeline render, 1200 s = 20 started minutes | $0.10 per started minute | $2.00 |
| Total for the audio and assembly | $2.125 |
Check before you publish
If the loop seam is audible, regenerate with a longer track, such as five minutes, and loop that instead. The generation price stays at $0.125.
Make a small library, one bed per class style, and reuse across the whole catalogue.
The hand-off between the two calls is the step people miss. When the music job completes, GET /v1/jobs/:id/result returns result.artifacts[] with an audio/mpeg file on media.sume.com. Timeline only accepts workspace media URLs, so import any outside file through POST /v1/media-imports first, and send the resulting URL as soundtrack.url. The Music docs also note that job.request.routed_model names the engine that actually ran, which is worth saving with the file.
- Listen to the loop point at least twice; a clean seam is the whole job here.
- Keep the voice cues and the bed at distinct levels; use
duck_dbif the teacher speaks. - Avoid rhythmic phrases in the prompt; any pulse becomes obvious when repeated.
- Confirm usage terms for generated music with Sume before selling class access.
Sources
Related posts
More in Media tools
- YouTube AI label and captions: burned-in subtitles on a real clip
YouTube's disclosure page exempts caption creation. Burning captions into your own footage with Sume is a $0.20 job for a clip up to 60 seconds.
- YouTube podcast thumbnail: 1:1, 10 MB on mobile, a square Sume still
YouTube podcast playlist thumbnails are 1:1, with a 10 MB limit on mobile and 50 MB on desktop. Make a square cover with the Sume images API at 1:1.
- YouTube upload bitrate: 8 Mbps at 1080p vs a 1080x1920 Short
YouTube recommends 8 Mbps for 1080p SDR (12 Mbps high frame rate). A 3-minute Short at 8 Mbps is about 180 MB. Sume trim conforms fps to 24, 25, 30 or 60.
- YouTube SDR uploads: BT.709 and what to check on a Sume MP4
YouTube recommends BT.709 for SDR, 4:2:0, progressive, moov atom up front. Sume trim exact re-encodes libx264 yuv420p; check color tags with ffprobe.
Written by Sume