Describe sounds in Short captions: burn [applause] cues with Sume

YouTube's caption help suggests text like [applause] or [thunder]. Burn those sound cues into a Short with Sume caption cues, no speech-to-text needed.

5 min readSume
All posts

To put sound descriptions such as [applause] or [thunder] on a Short, send them to Sume's video captions endpoint as cues. Each cue is a text plus start and end in seconds, and Sume burns exactly that text at those times without running speech-to-text. YouTube's help page suggests this kind of text for accessibility when you type captions by hand.

What YouTube's page says

YouTube's "Add subtitles & captions" page describes a type-manually method: you type or paste a transcript of your captions and subtitles and the timing is added automatically. For accessibility it suggests including descriptive text such as [applause] or [thunder] (page read 2026-10-05).

That text lives in a caption track that YouTube serves next to the video. A burned-in cue is different: it is part of the picture, so it shows on every platform that plays the file. Which one you want depends on where the Short will run. If you also upload the clip to TikTok or Reels, a burned cue travels with it.

The Sume side: cues instead of speech-to-text

Sume's video captions endpoint takes a public HTTPS video_url and either speech-to-text or authored text. The docs say you can send cues (or segments) with text, start and end, and that Sume then does not run speech-to-text. You can send only one of script_text, words, cues and segments in the same request.

This matters for sound descriptions because a doorbell or applause has no words for a transcript to find. If the clip is silent, speech-to-captions fails with caption_no_speech; authored cues are the documented way around it.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sound-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/short.mp4",
    "cues": [
      { "text": "[applause]", "start": 2.0, "end": 3.5 },
      { "text": "[thunder]", "start": 7.2, "end": 8.4 }
    ]
  }'

Keep the cue short and on screen long enough

A cue only helps if a viewer can read it. Keep the text to a few characters, give it a window of a second or more, and do not stack it on top of spoken captions in the same frame. Burned text also sits in the lower part of a vertical frame by default style rules, so check one still before you pay for the full batch.

  • Use square brackets so a sound description reads differently from speech.
  • One cue per sound, not one per second of the clip.
  • If the clip also has speech, burn the spoken captions in a second pass; the docs describe a restyle path that reuses word timings.
  • Read the price in the catalog before a batch. The docs quote $0.20 per accepted caption job for videos of at most 60 seconds, and say to check GET /v1/catalog for the live price.

When to leave it to YouTube

If a Short only runs on YouTube and you want viewers to be able to switch captions off, a typed caption track in Studio is the better fit. Sume's captions endpoint does not produce a caption file: the docs state it does not support SRT uploads, and the output is a captioned video.

Use burned cues when the text is part of the creative, when the same file is reused on other platforms, or when you want the sound descriptions styled like the rest of your captions.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume