YouTube Shorts tap to mute: sound-off captions from cues

A tap now mutes a Short. If your Sume clip has music only or a silent video, burn authored cues with video-captions so the message still reads.

5 min readSume
All posts

How do I caption a silent Short so it works when muted?

Send your own text as cues to Sume's video captions route. Each cue carries text, start and end in seconds, and Sume burns exactly that copy at those times without running speech-to-text, so it works on a clip with no speech at all.

This matters because YouTube's Shorts update lists simple muting through screen taps among the player changes, which makes sound-off viewing easier to reach.

Why a silent clip fails without cues

Speech-to-captions needs audible speech. According to the Sume docs, a silent clip fails as caption_no_speech with next_action: use_overlay_captions, which is a typed hint to pass authored cues instead. Check the clip first: a video inspect call with frames: false returns probe.has_audio without producing stills.

Which caption input fits which clip (Sume docs, read 2026-10-03)
SituationWhat to sendSpeech-to-text runs?
Clip with clear speech, wording is finevideo_url onlyYes
Clip with speech, you want your exact wordingvideo_url and script_textYes, used for timing
Silent clip or music-only bedvideo_url and cues (or segments)No
Same video, new looksource_caption_id and styleNo second run

A request that burns authored cues

Write one idea per cue, and make each cue last long enough to read. The docs say a standalone caption job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate, so confirm live pricing in the catalog before a bulk run. The source URL must be a fetchable public HTTPS video URL.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sound-off-cues-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/silent.mp4",
    "style": "slam",
    "cues": [
      {"text": "Three steps to a clean desk", "start": 0.2, "end": 2.5},
      {"text": "1. Clear it", "start": 2.5, "end": 4.5},
      {"text": "2. Wipe it", "start": 4.5, "end": 6.5}
    ]
  }'

Writing cues a muted viewer reads at a glance

Give each cue one idea, and let it stay on screen long enough that a person who looks away and back still catches it. Cues should not overlap in time. Start the first one in the first second so the hook is visible straight away, and end the last one before the clip does. Video inspect returns probe.duration_seconds, which tells you exactly how long the clip is before you write timings.

The route does not accept SRT uploads or provider task ids. The docs say to pass phrase-level text as cues or segments with text, start and end, so convert a subtitle file into that array in your own script before you send it.

Poll the job and read the result

The job runs asynchronously. Poll GET /v1/jobs/:id/status, then read the captioned video_url from the result or from GET /v1/video-captions/:id. To try another look on the same video, pass source_caption_id instead of video_url; the docs say that reuses the stored source and timings, and billing is unchanged because a restyle is still a render.

  • Do not mix script_text, words, cues and segments; the docs say they are mutually exclusive.
  • Start the first cue in the first second so a muted viewer gets the hook straight away.
  • Check one still from the output before you publish.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume