Captions for applause and music cues: author the cue text yourself

ElevenLabs Scribe can tag laughter and applause. On Sume, a silent clip returns caption_no_speech, so supply your own cues.

4 min readSume
All posts

If a clip has no speech, the Sume captions endpoint will not invent text: it returns caption_no_speech, and you pass your own cues or segments, for example "[applause]". ElevenLabs Scribe v2 can tag audio events like laughter and applause in a transcript, but that is the vendor's speech-to-text product and not a captions option on Sume.

So the answer to the search query is: write the cue text, give it timings, and let Sume render it in a caption style.

What the vendor offers

The ElevenLabs speech-to-text doc lists word-level timestamps, diarization for up to 32 speakers and audio tagging of events such as laughter and applause for Scribe v2. That is useful when you want an accessibility transcript with sound descriptions generated automatically.

What Sume captions do

The video captions endpoint charges $0.20 for videos up to 60 seconds. It transcribes the audio unless you skip transcription by supplying words, cues or segments, which are mutually exclusive with script_text. A silent clip with nothing to transcribe returns caption_no_speech, and the docs point to cues or segments as the way to caption it.

That behavior is a feature for sound labels: you decide which sounds deserve a caption, what the text says, and when it appears. You control the wording, so a bracketed label like [applause] or [upbeat music] appears exactly as written.

  • Cues carry the text and the start and end times you choose.
  • Pick a style from the docs; Latin text defaults to slam when you omit a style.
  • Use design overrides for placement and phrasing if the label sits over a busy area.
  • Estimated duration over 60 seconds is rejected.

Mixing speech and labels

A clip with a speaker and a laughing crowd needs two sources. Transcribe the speech by letting the endpoint run, then restyle or re-render with your own cues that include the label lines, or build the full cue list yourself from a transcript you trust. Because words, cues and segments skip transcription, you can assemble spoken lines and sound labels into one list and send it once.

A restyle with source_caption_id re-uses an earlier caption without transcribing again, which helps if you only want to change the look.

Sound-label captions: vendor transcript tags versus Sume cues (read 2026-10-03)
NeedElevenLabs Scribe v2Sume video captions
Spoken wordsTranscribed, word-level timestampsTranscribed unless you supply words, cues or segments
Laughter or applauseAudio tagging in the transcriptYou write the cue text and timings
Silent clipNot the focus of the doccaption_no_speech; pass cues or segments

Style and legibility

Sound labels should be short and visually quieter than speech. Use a smaller placement or a plainer style, and keep brackets so viewers can tell a label from dialogue. Check the result on the target platform's UI safe areas before you publish.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume