Cut AI clips to the beat: BPM to Timeline start times in Python
Ask Lyria for a tempo, turn it into cut times, and render the clips on the beat with Sume Timeline. Python sketch, cost, and how to check the tempo.

To cut clips on the beat, ask the music prompt for a tempo as a number, compute one cut length from it (60 divided by BPM, times beats per cut), and set each Timeline slot's start to a multiple of that length. At 120 BPM, one cut per bar is 2.0 seconds, so eight clips fill 16 seconds.
The tempo in the prompt is a creative direction, not a setting. Sume Music has no tempo parameter, so listen to the track and fix the offset before you render.
How do I turn a tempo into cut times?
Music 1.0 and the Music Router accept a prompt, an optional image and no duration or tempo field. The docs list tempo as a number (for example "72 BPM") as one axis of the brief. Google's Gemini API changelog lists Lyria 3.5 as generally available on 2026-09-03 with fine-grained duration and structure control, and Sume's Music Router runs it today.
Beat length is 60 / BPM seconds. Multiply by the beats you want per cut.
| BPM | Seconds per beat | Seconds per bar |
|---|---|---|
| 90 | 0.667 | 2.667 |
| 100 | 0.600 | 2.400 |
| 120 | 0.500 | 2.000 |
| 128 | 0.469 | 1.875 |
| 140 | 0.429 | 1.714 |
What does the Timeline request look like?
Timeline 1.0 takes one audio spine and ordered video slots. Slot starts must increase, the first start must be 0, and each slot must be at least 0.2 seconds. Use the music file as the spine, so the cuts and the track share one clock.
The sketch below prints the slot list for eight clips on a 120 BPM track. Replace the URLs with your own media.sume.com artifacts.
import json, math
bpm = 120 # the tempo you asked for in the music prompt
beats_per_cut = 4 # one cut per bar
first_downbeat = 0.0 # seconds; set after you listen
clips = [f"https://media.sume.com/artifacts/artf_demo/clip{i}.mp4" for i in range(8)]
cut = 60 / bpm * beats_per_cut
total = cut * len(clips)
body = {
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/track.mp3",
"duration_seconds": math.ceil(total),
"source_in": first_downbeat},
"video": [{"source_url": u, "start": round(i * cut, 3), "duration": round(cut, 3)}
for i, u in enumerate(clips)],
}
print(json.dumps(body["video"][:3], indent=1))
print("cut every", cut, "s; total", total, "s; billable minutes", math.ceil(total / 60))
How do I check that the track really sits on the grid?
Lyria can drift from the tempo you asked for, and the first downbeat may not land at 0:00. Play the track against a stopwatch, or tap along for ten seconds, and note the real first downbeat. Put that value in audio.source_in: the docs describe it as the in-point into a single audio spine, and the output length is still duration_seconds.
If cuts drift against the music after 20 or 30 seconds, the real tempo differs from your number. Re-derive the cut length from what you hear, or generate another take. The same prompt can give a different track each time, so keep the artifact you accepted.
What does this cost?
Music is a fixed $0.125 per accepted generation in the Music 1.0 docs, and the Music Router charges the same fixed price whichever engine runs. Timeline charges $0.10 per output minute, rounded up, so a 16-second render is one billable minute. POST /v1/timeline-1.0/plan is an unbilled preflight that returns the billable minutes and estimated cost, so run it before the render.
A 16-second beat-cut edit is $0.125 plus $0.10, plus whatever the eight clips cost. Do not claim the cuts are musically perfect: check the render with your eyes and ears.
Sources
Related posts
More in Developers
- Decart lucy-latest vs a pinned Lucy model; Sume catalog ids
Decart's lucy-latest alias can move while legacy Lucy Clip costs $0.15 per second against $0.04 for Lucy 2.5. Why pin a model id, and how to do it on Sume.
- Detect new AI image models: diff Sume GET /v1/images/models
Image models arrive weekly. A short Python diff against GET /v1/images/models tells you when Sume adds or retires an image model id, with no news feed to watch.
- Download a finished Seedance or Kling MP4 from the Sume API
Two ways to get the file: GET /v1/videos/{id}/content with your key, or the hosted artifact URL from /v1/jobs/{id}/result. A Python download script.
- Draft at 1K, finish at 2K: read the image resolution descriptor
Sume's resolution tiers run 512 to 4K, but each model lists its own. A script that reads the descriptor, then a draft-then-final pattern for ad images.
Written by Sume