Turn LRC lyrics into caption cues for a lyric video (Python)

A 26-line Python script converts an LRC lyrics file into the cues array for POST /v1/video-captions, so a silent lyric video is captioned for $0.20.

4 min readSume
All posts

If you have lyrics with timestamps in LRC format, you can burn them onto a video by converting each line to a cues entry with text, start and end. The video captions docs say authored cues skip speech-to-text and burn your text at exactly those times, which is the path for silent clips. One caption job is $0.20.

LRC lines look like [00:01.50] Silent night: minutes, seconds, then the words. A line ends when the next one starts, so the converter pairs each line with the next timestamp. An empty final line, like the last one in the sample, marks the end of the last lyric.

The script

It parses the timestamps, builds the cues and prints the request body for POST /v1/video-captions. Replace the sample URL with a video hosted on your workspace.

import json, re

LRC = """[00:01.50] Silent night
[00:05.00] Holy night
[00:09.25] All is calm
[00:13.00]"""

def parse(text):
    rows = []
    for line in text.splitlines():
        m = re.match(r"\[(\d+):(\d+(?:\.\d+)?)\]\s*(.*)", line)
        if m:
            t = int(m.group(1)) * 60 + float(m.group(2))
            rows.append((t, m.group(3).strip()))
    return rows

def cues(rows):
    out = []
    for (start, words), (end, _) in zip(rows, rows[1:]):
        if words:
            out.append({"text": words, "start": round(start, 2), "end": round(end, 2)})
    return out

body = {"video_url": "https://media.sume.com/artifacts/example/lyrics.mp4",
        "style": "punch", "cues": cues(parse(LRC))}
print(json.dumps(body, indent=2))

Using the output

Send the printed JSON with your API key and an Idempotency-Key header. Only one of script_text, words, cues and segments may be sent in a request, so do not combine them. Pick a style from the docs list, such as punch or tiktok-green.

The caption price covers videos of up to 60 seconds. A full-length song video needs to be split into parts, with one cue list for each part and times counted from the start of that part.

Edge cases

  • Lines with no text before the next timestamp are skipped, so instrumental gaps stay clear of captions.
  • Times are rounded to two decimals in seconds.
  • Check the first and last cue on a test render before you run the full batch.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume