AI song captions: Lyria lyrics are metadata, so burn your own cues

result.lyrics on a Sume music job is model-reported, not measured. For a lyric video, burn captions with your own cues: music $0.125 plus a $0.20 caption job.

5 min readSume
All posts

Do not time a lyric video from result.lyrics on a Sume music job. The docs call those lyrics model-reported metadata, not an audio measurement. Write the lines and their start and end seconds yourself and send them to POST /v1/video-captions as cues, which skips speech-to-text. A 30-second lyric clip is then $0.125 for the music plus $0.20 for the caption job, $0.325 before any render.

What the docs say about lyrics

The Music Router docs say that when they are present, result.lyrics carries the model-reported lyrics or section map. The Music 1.0 docs add that provider lyrics can describe tempo and structure but are not a measurement of the audio. Section markers in the prompt, such as [0:00-0:30] Intro: ..., are a request to the model, not a promise that the audio lands on those seconds.

Why not let speech recognition do it

Standalone captions run speech-to-text unless you send words, cues or segments, and speech-to-captions works only when the clip has audible speech; otherwise the job fails with caption_no_speech. The docs do not promise accurate recognition of sung vocals, so for a lyric video the safe route is authored timings. You listen once, mark where each line starts and ends, and send them.

Three ways to caption a 30-second AI song, rates read 2026-10-09
RouteWhat sets the timingFails whenCost of the caption job
Speech-to-text (no cues)Recognition of the audioClip has no audible speech: caption_no_speech$0.20
script_textSpeech-to-text word times, text from your scriptAlignment fails: script_alignment_mismatch$0.20
cuesYour own start and end secondsNever for timing; text must suit the style$0.20

The request

Each cue is text, start and end in seconds. You may send only one of script_text, words, cues and segments. Pick a style that fits the language, and for Korean lyrics use a Hangul style rather than a Latin one, because Latin styles are rejected for Korean text with caption_hangul_text_latin_style.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: lyric-clip-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/song-clip.mp4",
    "style": "slam",
    "cues": [
      { "text": "Streetlights on the water", "start": 1.2, "end": 4.0 },
      { "text": "Nobody left to call", "start": 4.2, "end": 7.1 }
    ]
  }'

Marking cues without guessing

Play the finished song and write down, in seconds, where each line starts and where its last word ends. Leave a short gap between cues so two lines never overlap on screen, and keep each cue to a phrase you can read in the time it is shown. A cue is overlay text, so a repeated chorus is just the same text with new times.

If the video has no vocals at all, say so in the prompt, because the docs advise ending with the clause "Instrumental, no vocals." A silent-vocal clip captioned by recognition fails with caption_no_speech and its next_action points back to overlay captions, which is the cues route.

Practical order

Caption after the cut, not before.

  • Generate the track and measure it; length is steered by the prompt, not a parameter.
  • Join the track to the picture in a Timeline render ($0.10 per output minute).
  • Caption the finished MP4 last, so the burned text is never cut by a later trim.
  • Keep your cues in version control; changing a style later with source_caption_id does not need new timings.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume