Toy store spot cut to the beat: a Lyria track as the Timeline spine
Make a 15-second toy ad with the music as the spine: Lyria 3.5 track, audio.source_in to start at the best bar, then slots cut on the beat for $0.225 on Sume.

To cut a toy ad to the beat, make the music the Timeline spine: send the Lyria track as audio.url, set audio.source_in to the second where the section you want starts, and set audio.duration_seconds to 15. The render is one minute at $0.10 and the track is $0.125, so the spot costs $0.225.
Most Sume timelines put a voice-over on the spine and the music in soundtrack. A toy ad has no narrator. Using the track itself as the spine means the length, the in-point and the cut points are all tied to the music you picked.
A spine that is music
audio.source_in sets the in-point into a single url spine, and it is not permitted with parts. The output length is still duration_seconds, so a 60-second track with source_in: 24 and duration_seconds: 15 plays seconds 24 to 39. audio.gain_db lifts or lowers the spine between -60 and 12.
The slots then carry the beat. With a 120 BPM track a beat is 0.5 seconds, so slots of 1, 1.5 or 2 seconds land on beats if the in-point itself lands on one; you hear the track and read the beat from the audio, nothing in the API does it for you. Slots are at least 0.2 seconds, and declared starts are authoritative, so the cuts sit exactly where you place them.
A good first pass is to take the section that begins right after the track's intro, because it usually has the clearest pulse. Download the audio, listen for the first strong downbeat, and use that time as source_in; if the cuts feel half a beat late, move source_in by a fraction of a second rather than editing every slot.
| Slot | start | duration | Beats |
|---|---|---|---|
| Hero toy, box | 0 | 2 | 4 |
| Close-up, moving part | 2 | 1.5 | 3 |
| Child's hands, play | 3.5 | 2.5 | 5 |
| Three toys together | 6 | 3 | 6 |
| Pack shot and name | 9 | 6 | 12 |
Picking the section
Ask Lyria for what you need in the prompt: a 60-second track is long enough to audition sections. Because Lyria has no duration field, write the length into the text, and put exclusions in the positive prompt; for example "playful xylophone and hand claps, 120 BPM, bright, steady beat, no vocals". The router also rejects negative_prompt, so say what to avoid inside the prompt.
Check the in-point by ear on the downloaded track and then plan the render. The plan call is unbilled and returns duration_seconds, segment_count and billable_minutes. It cannot predict short-source pad or loop warnings, so a clip shorter than its slot is only caught on the real render.
Spine from the track
import os
import uuid
import requests
starts = [(0, 2), (2, 1.5), (3.5, 2.5), (6, 3), (9, 6)]
clips = [os.environ[f"TOY_{i}"] for i in range(5)] # media.sume.com URLs
video = [{"source_url": u, "start": s, "duration": d}
for u, (s, d) in zip(clips, starts)]
body = {
"audio": {"url": os.environ["TRACK_URL"], "source_in": 24,
"duration_seconds": 15, "gain_db": -3},
"video": video,
"output": {"fade_out_seconds": 1},
}
r = requests.post("https://api.sume.com/v1/timeline-1.0/render", json=body,
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"toy-{uuid.uuid4()}"},
timeout=60)
print(r.status_code, r.json().get("request_id"))
Fades, URLs and trying three tracks
Output edge fades are output.fade_in_seconds and output.fade_out_seconds (0 to 5 each, together within the output length), so fade_out_seconds: 1 ends picture and sound on the last beat. The last slot covers to 15 seconds, within the 0.5-second coverage rule.
A few practical notes. First, every URL must be this workspace's media.sume.com artifact or asset, so import footage and the track first. Second, to try a different track you change one URL and source_in, not the cut list, which makes it a cheap way to test three feels.
Three tracks at $0.125 each and three renders at $0.10 is $0.675 for three versions of the same spot. For beds under a voice, see music beds and ducking under a voice. Keep any age or safety statements for the pack, not for the generated frame.
If a toy clip comes from an AI video model rather than your camera, check it before you cut to it: a moving part that appears to do something the real toy cannot is a mismatch between the ad and the product, and the cut list will not catch it. Real footage of the real toy is the safest source for the close-ups.
Sources
Related posts
More in Use cases
- Training knowledge-check clips with an AI avatar, one per question
Make a question clip and an answer clip per quiz item with one Sume avatar: 10-second scripts, a silence beat for thinking time, and the cost for 20 questions.
- Transcribe a 30-minute video: three 10-minute detaches and STT, $0.33
A 30-minute file is too long for one detach or one STT call. Cut three 10-minute audio ranges at $0.01 each, then transcribe each at $0.01 per minute: $0.33.
- Translate a product label in place into 4 languages with Ideogram 4.5
Swap approved label copy per market with one Ideogram 4.5 edit each on Sume: a copy-deck prompt, $0.15 for four languages at low, and a proofing list.
- Translate chart and diagram labels in an image with Ideogram 4.5
Swap the axis titles, legend and callouts of a chart image into another language with one Ideogram 4.5 edit per language: $0.0375 each at low on Sume.
Written by Sume