Captions for a video with loud background music: script text or cues

Music can bury speech and trip speech-to-text. Sume documents three ways to supply your own wording: script_text, a words array or cues. Which to pick.

5 min readSume
All posts

If a video has loud background music and the automatic captions miss words, stop relying on speech recognition for the wording and supply your own text. Sume's caption endpoint accepts script_text, a words array, or cues, and each of those makes the wording yours. Only script_text keeps the speech-to-text timings; words and cues skip speech-to-text entirely.

Sume does not publish a noise-suppression step or an accuracy number for music-heavy audio, so this post makes no such claim. It is about what you control.

Option 1: script_text, STT timings

Send the finished script as script_text with the video_url. Per the caption docs, Sume keeps the speech-to-text word timings as the source of truth for time and aligns the burned-in text to your script. You get your spelling and the recognizer's timing. The alignment can fail with script_alignment_mismatch or script_alignment_failed, and the documented next action is to simplify the script or omit it. If speech is so buried that timing is also unreliable, this path may fail.

The limit is 8,000 characters for script_text.

Option 2: words or cues, no recognizer

If you have timings from elsewhere, for example the text-to-speech job that made the voiceover, send words (single tokens with text, start, end) or cues (phrase cards up to 400 characters each, up to 200 per request). With these Sume does not run speech-to-text and burns the text at the times you give. You can send only one of script_text, words, cues and segments in a request.

This is the dependable route for a voiceover you generated yourself, because you already know both the words and when they were spoken.

Wording sources on a standalone caption job, from the Sume docs and request schema read 2026-10-07
You sendWho supplies wordingWho supplies timingSpeech-to-text runs
Nothing but video_urlRecognizerRecognizerYes
script_text (max 8,000 characters)YouRecognizerYes
words (max 1,200 tokens)YouYouNo
cues or segments (max 200)YouYouNo

Option 3: lower the music before you caption

If you control the mix, caption the voice-only version. If you built the video yourself, caption the version before the music is added, or send the voiceover's own word times as words. If you only have the finished file, audio detach can pull the track as a 16 kHz mono WAV for a transcription check, but it does not remove music, and the ffmpeg fields are rejected so you cannot add a filter.

Price

A standalone caption job is $0.20 for a video of up to 60 seconds, whichever wording source you use. The caption price does not change with the source. A transcription probe of the audio is separate: $0.01 per audio minute, so a 60 second clip checked with video inspect costs one cent plus the inspect compute.

A sensible order is: try the default once, read the output, and if words are wrong, re-burn with script_text. You can restyle an existing caption by source_caption_id without running speech-to-text a second time, and words can correct a single word.

Checks before you ship

Play the burned result with the sound on, because mistimed words are easier to hear than to see. Look at the first and last caption: if the first appears before the voice starts, the recognizer picked up music as speech. Scan names, numbers and product terms; those are where a recognizer under loud music fails first, and where script_text helps most.

If the clip has no spoken words at all, for example an instrumental montage, the standard job fails with caption_no_speech and the documented fix is authored cues. See captions on a silent AI clip for that case.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume