Lyric video captions for an AI song: pass the lyrics as script_text

Lyria 3.5 returns lyrics with the audio. Sume's captions API can take them as script_text to caption a video that carries the song. Steps and limits.

4 min readSume
All posts

For lyric video captions, generate the song, render it into a video, and call the Sume captions API with the lyrics as script_text. Google's Gemini post, read 2026-10-01, describes Lyria 3.5 for tracks with vocals and lyrics; Sume's Music Router returns them in result.lyrics when the model reports them.

What are the steps?

Generate with POST /v1/music-router/generate. Render an MP4 with Timeline from the audio and a still. Then POST /v1/video-captions with video_url, a style, and script_text. Captions cost $0.20 per job for videos up to 60 seconds.

What are the limits?

result.lyrics is the model's report, not a transcript of the audio. Sume captions need audible speech; the docs name caption_no_speech for silent clips. Sung words may not be treated like speech, so test a sample, and if it fails supply cues with text and timing yourself.

Can I keep my styling?

Yes: style, font and design overrides exist, and source_caption_id restyles an earlier caption. There is no SRT upload.

Which endpoint does each step use?

The steps above, in one place.

Lyric caption pipeline, read 2026-10-01.
StepEndpoint or fieldNote
Make the songPOST /v1/music-router/generateLyrics in result.lyrics when reported
Render videoTimelineAudio plus a still, MP4 output
CaptionPOST /v1/video-captionsvideo_url, style, script_text
Caption price$0.20 per jobVideos up to 60 seconds

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume