Turn LRC lyrics into caption cues for a lyric video (Python)
A 26-line Python script converts an LRC lyrics file into the cues array for POST /v1/video-captions, so a silent lyric video is captioned for $0.20.

If you have lyrics with timestamps in LRC format, you can burn them onto a video by converting each line to a cues entry with text, start and end. The video captions docs say authored cues skip speech-to-text and burn your text at exactly those times, which is the path for silent clips. One caption job is $0.20.
LRC lines look like [00:01.50] Silent night: minutes, seconds, then the words. A line ends when the next one starts, so the converter pairs each line with the next timestamp. An empty final line, like the last one in the sample, marks the end of the last lyric.
The script
It parses the timestamps, builds the cues and prints the request body for POST /v1/video-captions. Replace the sample URL with a video hosted on your workspace.
import json, re
LRC = """[00:01.50] Silent night
[00:05.00] Holy night
[00:09.25] All is calm
[00:13.00]"""
def parse(text):
rows = []
for line in text.splitlines():
m = re.match(r"\[(\d+):(\d+(?:\.\d+)?)\]\s*(.*)", line)
if m:
t = int(m.group(1)) * 60 + float(m.group(2))
rows.append((t, m.group(3).strip()))
return rows
def cues(rows):
out = []
for (start, words), (end, _) in zip(rows, rows[1:]):
if words:
out.append({"text": words, "start": round(start, 2), "end": round(end, 2)})
return out
body = {"video_url": "https://media.sume.com/artifacts/example/lyrics.mp4",
"style": "punch", "cues": cues(parse(LRC))}
print(json.dumps(body, indent=2))Using the output
Send the printed JSON with your API key and an Idempotency-Key header. Only one of script_text, words, cues and segments may be sent in a request, so do not combine them. Pick a style from the docs list, such as punch or tiktok-green.
The caption price covers videos of up to 60 seconds. A full-length song video needs to be split into parts, with one cue list for each part and times counted from the start of that part.
Edge cases
- Lines with no text before the next timestamp are skipped, so instrumental gaps stay clear of captions.
- Times are rounded to two decimals in seconds.
- Check the first and last cue on a test render before you run the full batch.
Sources
Related posts
More in Use cases
- Lyria 3.5 screens prompts: brand slogans and lyrics in an ad song
Google says every Lyria prompt is safety-screened and it cannot output copyrighted lyrics or named artist voices. How to write a brand jingle brief that passes.
- Make 3-minute hold music for a voice agent
ElevenLabs agents can play hold audio up to 180 s and 40 MB. Generate a short loop with Sume music generation and join it into 180 s with timeline audio.
- Meta ad video length by placement: 15 s to 240 min
Meta's placement chart sets different lengths and ratios per placement, from 15 s Messenger Stories to 240 min Feed. A render plan by placement.
- Meta's AI info label: signals or self-disclosure, and a trimmed clip
Meta applies its AI info label from industry-standard signals or self-disclosure. A trim or re-encode may drop file metadata, so disclose at upload too.
Written by Sume