Captions for applause and music cues: author the cue text yourself
ElevenLabs Scribe can tag laughter and applause. On Sume, a silent clip returns caption_no_speech, so supply your own cues.

If a clip has no speech, the Sume captions endpoint will not invent text: it returns caption_no_speech, and you pass your own cues or segments, for example "[applause]". ElevenLabs Scribe v2 can tag audio events like laughter and applause in a transcript, but that is the vendor's speech-to-text product and not a captions option on Sume.
So the answer to the search query is: write the cue text, give it timings, and let Sume render it in a caption style.
What the vendor offers
The ElevenLabs speech-to-text doc lists word-level timestamps, diarization for up to 32 speakers and audio tagging of events such as laughter and applause for Scribe v2. That is useful when you want an accessibility transcript with sound descriptions generated automatically.
What Sume captions do
The video captions endpoint charges $0.20 for videos up to 60 seconds. It transcribes the audio unless you skip transcription by supplying words, cues or segments, which are mutually exclusive with script_text. A silent clip with nothing to transcribe returns caption_no_speech, and the docs point to cues or segments as the way to caption it.
That behavior is a feature for sound labels: you decide which sounds deserve a caption, what the text says, and when it appears. You control the wording, so a bracketed label like [applause] or [upbeat music] appears exactly as written.
- Cues carry the text and the start and end times you choose.
- Pick a style from the docs; Latin text defaults to slam when you omit a style.
- Use design overrides for placement and phrasing if the label sits over a busy area.
- Estimated duration over 60 seconds is rejected.
Mixing speech and labels
A clip with a speaker and a laughing crowd needs two sources. Transcribe the speech by letting the endpoint run, then restyle or re-render with your own cues that include the label lines, or build the full cue list yourself from a transcript you trust. Because words, cues and segments skip transcription, you can assemble spoken lines and sound labels into one list and send it once.
A restyle with source_caption_id re-uses an earlier caption without transcribing again, which helps if you only want to change the look.
| Need | ElevenLabs Scribe v2 | Sume video captions |
|---|---|---|
| Spoken words | Transcribed, word-level timestamps | Transcribed unless you supply words, cues or segments |
| Laughter or applause | Audio tagging in the transcript | You write the cue text and timings |
| Silent clip | Not the focus of the doc | caption_no_speech; pass cues or segments |
Style and legibility
Sound labels should be short and visually quieter than speech. Use a smaller placement or a plainer style, and keep brackets so viewers can tell a label from dialogue. Check the result on the target platform's UI safe areas before you publish.
Sources
Related posts
More in Media tools
- Check a product-swap video edit for the old product with frames
After a prompted product swap on a video, pull matching stills from the source and the edit and compare them to catch the old product. Sume docs for each step.
- Check a video crop before you pay: free video-filter /check
Validate a crop or dim program with POST /v1/video-filter/check, which is free, then submit the $0.02 encode. Diagnostics, refusal codes and next_action.
- Check the transcript before captions: Scribe v2 error tiers
ElevenLabs lists Scribe v2 error rates by language group. Before burning captions on Sume, pass script_text or your own words so a wrong transcript never ships.
- Children's captions at 17 characters per second: Netflix rule and Sume
Netflix's English style guide sets 20 chars per second for adults and 17 for children. See how to approximate it with Sume caption phrasing overrides.
Written by Sume