AI workout music for fitness videos: tempo and intervals in a prompt
Make AI workout music by writing tempo as a number and the interval plan as timestamps. Sume has no BPM field, so here is how to prompt, check and trim a track.

To get AI workout music with a given tempo, write the tempo into the prompt as a number, and write the interval plan as timestamped sections. Sume's Music Router has no BPM, key or duration field; the Music docs say to steer tempo and length in the prompt text, for example "142 BPM half-time" or [0:00-0:30] Intro: .... The result is a direction, and the docs tell you to verify the audio rather than assume it.
A workout track has a job a background bed does not: it has to change when the exercise changes. If your video runs 40 seconds on and 20 off, the music should lift on the first and ease on the second. This post shows a prompt that does that, how to check the tempo it gave you, and how to cut it to length.
How do I write intervals as sections?
The docs list "Tempo as a number" as one of seven brief axes, with examples such as "72 BPM" and "142 BPM half-time". They also show timestamped sections inside a prompt. Put them together, and each interval becomes a section with its own energy. The values below are examples to adapt, not figures from the docs.
| Section | Example prompt line | What changes |
|---|---|---|
| [0:00-0:20] Warm-up | Steady pulse at 100 BPM, soft kick, filtered synth | Low energy, room to talk |
| [0:20-0:50] Work | Same tempo, full drums, driving bass, lead enters | Energy up, a clear downbeat |
| [0:50-1:10] Rest | Drums drop out, pad and light hats | Breathing space |
| [1:10-1:40] Work | Return of full drums, brighter lead | Second push |
| [1:40-2:00] Cool-down | Slow to a soft pad, no drums | Clean end |
How do I check the tempo I got?
There is no tempo readback in the result. result.lyrics can carry a model-reported tempo and structure, but the docs call that metadata, not an audio measurement. To check a take, tap along against a stopwatch for one section, or load the file in any editor that shows a beat grid.
Because there is no seed, the same prompt gives a different take each time, so a miss is cheap to retry. Each accepted generation is billed at the fixed Music price. Change the tempo number or one section line per attempt and keep the rest, so you learn what moved the result.
What does the request look like?
A sample brief, with the plan above and the length stated in the prompt:
curl -X POST https://api.sume.com/v1/music-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: hiit-track-001" \
-d '{
"prompt": "Driving electronic workout track, 128 BPM, A minor. Four-on-the-floor kick, rolling bass, bright saw lead. [0:00-0:20] Warm-up: filtered pulse. [0:20-0:50] Work: full drums, lead enters. [0:50-1:10] Rest: drums drop, pad only. [1:10-1:40] Work: full drums return, brighter. [1:40-2:00] Cool-down: slow pad. A 2-minute track. Instrumental, no vocals."
}'What if the video is longer or shorter than the track?
Google's Lyria page says Lyria 3.5 makes tracks from a 60-second clip up to three minutes, so a two-minute plan fits. A video longer than the track needs a different approach. Sume's Timeline soundtrack can loop a bed, but a loop restarts the interval plan, which only works if your workout repeats the same cycle.
If the take runs long, trim it rather than regenerate. Timeline audio can split a file by ranges with a start and an optional end, and returns a new file with its own audio_url. Output is wav by default, so a trimmed workout cut stays sample-exact, and the job is a flat $0.01. Ask for a soft, sustained final section if you plan to cut early, because Timeline audio has no fade option.
- Longer video, same repeating cycle: use
soundtrack.loopin a Timeline render. - Longer video, changing plan: write a longer prompt, up to a few minutes, or join two tracks with Timeline audio
concat. - Shorter video: split with a range such as start 0, end 90, and let the video's own fade handle the end.
- Cut to the beat: see the stored guide on BPM to Timeline start times.
How do I fit a coach's voice over it?
If a coach talks over the track, make the voice the audio spine and the music the soundtrack. A Timeline render then ducks the music by duck_db, from 0 to 20, while the voice speaks, and gain_db sets the bed level; if you omit it, Sume's compiler applies minus 16 dB. The stored ducking guide has the request body. Start with a modest duck and raise it only if cues are lost, because a heavy duck on a track with sharp section changes makes the lift between intervals less obvious.
What will the prompt not guarantee?
Two honest limits. Sume will not hold an exact BPM, and it will not keep a section to the second you named, so the cues in your video may drift a few seconds from the music. Plan the edit around what you got back, not around what you asked for. And AI vocals under a coach's voice usually get in the way, so the closing clause "Instrumental, no vocals" is worth keeping.
Sources
Related posts
More in Use cases
- Animate a locally generated image with Sume video: first frame
Made a still with a local open-weights model? Host it at a public HTTPS URL, send it as first_frame to POST /v1/videos, and poll. Plus the licence check first.
- Can a brand make the video for a creator's Meta partnership ad?
Meta's branded content policy says creators cannot be paid to post content they were not involved in making. What that means for brand-made AI clips.
- Can I make an AI child character for TikTok? The under-18 rule
TikTok does not allow AIGC with the likeness of anyone under 18. What that says about child-like avatars, and what Sume's avatar tools record.
- Can viewers see less AI video on TikTok? The AIGC control
TikTok is testing a Manage Topics control for how much AI-generated content shows in For You. What TikTok says and what it means for AI video creators.
Written by Sume