AI song captions: Lyria lyrics are metadata, so burn your own cues
result.lyrics on a Sume music job is model-reported, not measured. For a lyric video, burn captions with your own cues: music $0.125 plus a $0.20 caption job.

Do not time a lyric video from result.lyrics on a Sume music job. The docs call those lyrics model-reported metadata, not an audio measurement. Write the lines and their start and end seconds yourself and send them to POST /v1/video-captions as cues, which skips speech-to-text. A 30-second lyric clip is then $0.125 for the music plus $0.20 for the caption job, $0.325 before any render.
What the docs say about lyrics
The Music Router docs say that when they are present, result.lyrics carries the model-reported lyrics or section map. The Music 1.0 docs add that provider lyrics can describe tempo and structure but are not a measurement of the audio. Section markers in the prompt, such as [0:00-0:30] Intro: ..., are a request to the model, not a promise that the audio lands on those seconds.
Why not let speech recognition do it
Standalone captions run speech-to-text unless you send words, cues or segments, and speech-to-captions works only when the clip has audible speech; otherwise the job fails with caption_no_speech. The docs do not promise accurate recognition of sung vocals, so for a lyric video the safe route is authored timings. You listen once, mark where each line starts and ends, and send them.
| Route | What sets the timing | Fails when | Cost of the caption job |
|---|---|---|---|
| Speech-to-text (no cues) | Recognition of the audio | Clip has no audible speech: caption_no_speech | $0.20 |
| script_text | Speech-to-text word times, text from your script | Alignment fails: script_alignment_mismatch | $0.20 |
| cues | Your own start and end seconds | Never for timing; text must suit the style | $0.20 |
The request
Each cue is text, start and end in seconds. You may send only one of script_text, words, cues and segments. Pick a style that fits the language, and for Korean lyrics use a Hangul style rather than a Latin one, because Latin styles are rejected for Korean text with caption_hangul_text_latin_style.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lyric-clip-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/song-clip.mp4",
"style": "slam",
"cues": [
{ "text": "Streetlights on the water", "start": 1.2, "end": 4.0 },
{ "text": "Nobody left to call", "start": 4.2, "end": 7.1 }
]
}'Marking cues without guessing
Play the finished song and write down, in seconds, where each line starts and where its last word ends. Leave a short gap between cues so two lines never overlap on screen, and keep each cue to a phrase you can read in the time it is shown. A cue is overlay text, so a repeated chorus is just the same text with new times.
If the video has no vocals at all, say so in the prompt, because the docs advise ending with the clause "Instrumental, no vocals." A silent-vocal clip captioned by recognition fails with caption_no_speech and its next_action points back to overlay captions, which is the cues route.
Practical order
Caption after the cut, not before.
- Generate the track and measure it; length is steered by the prompt, not a parameter.
- Join the track to the picture in a Timeline render ($0.10 per output minute).
- Caption the finished MP4 last, so the burned text is never cut by a later trim.
- Keep your cues in version control; changing a style later with
source_caption_iddoes not need new timings.
Sources
Related posts
More in Media tools
- Meta's 4 GB video cap at Sume's 30-minute render limit: 17.8 Mbps
Meta ad video tops out at 4 GB. A 30-minute Sume Timeline render can average up to 17.8 Mbps and fit; a 15-minute Reels ad allows 35.6 Mbps. Table inside.
- Meta lists Reels ads from 0 seconds; Sume's shortest trim is 0.2
Meta's Reels ad page lists lengths from 0 seconds, Feed from 1. Sume's shortest trim is 0.2 s and shortest Timeline render 1 s; shortest clip per model.
- Kling 3.0 Motion Control API: 7.2 s driver video bills 8 s, $1.26
Sume's Kling 3.0 Motion Control costs $0.1575 per second rounded up from the driving video, 1 to 30 seconds. Fields, orientation, and the cost table.
- A 12-second sting scored from a product still: image_url on Sume music
Sume music takes an optional image_url and no duration field. Put "a 12-second track" in the prompt; the job is a flat $0.125 whatever the image or the length.
Written by Sume