YouTube automatic captions not available: what to do instead

YouTube lists silence, overlapping speakers and poor audio as reasons auto captions fail. Burn your own with Sume cues or a script_text-aligned caption job.

5 min readSume
All posts

When YouTube says automatic captions are not available for a video, its own help page gives the causes: audio still processing, an unsupported language, a video over the length limit, poor sound quality or unrecognised speech, a long silence at the start, or several overlapping speakers or languages. You can only fix the first with time. For the rest, a practical route is to burn captions into the file yourself, and Sume's video captions endpoint can do that from your own text and timings.

The YouTube facts here come from Use automatic captioning, read on 2026-10-02. The Sume facts come from the video captions and video inspect docs. Burned-in captions are pixels in the video; they do not populate YouTube's own caption track, so viewers cannot toggle them off, and this post does not claim otherwise.

Why does YouTube skip automatic captions on some videos?

The help page says automatic captions are generated by machine learning, so quality varies, and that mispronunciations, accents, dialects or background noise can make them misrepresent what was said. It also lists the conditions under which captions are not generated at all. Two of them matter most for edited footage: extended silence at the start, which is common in b-roll openers and cinematic intros, and multiple overlapping speakers or simultaneous languages, which is common in panels and reaction videos.

YouTube's advice is to review automatic captions in YouTube Studio under Subtitles and edit the parts that were not transcribed properly. That works when captions exist. When none were produced, you need another source of text and timing.

Which Sume caption input fits each failure?

Sending text fields together is an error: script_text, words, cues and segments are mutually exclusive. A silent clip sent without cues fails as caption_no_speech with next_action: use_overlay_captions, which is the same gap YouTube describes.

Caption inputs for a YouTube auto-caption gap, read 2026-10-02
Cause on YouTube's listSume caption inputWhat happens
Silent clip or long silent opener, no speech at allcues (or segments) with text, start, endBurns your authored phrases at your times, no speech-to-text
Speech present, you have the scriptscript_textSpeech-to-text timings stay the source of truth; wording aligns to your script
Speech present, no scriptOmit all text fieldsSpeech-to-text wording is burned in
You already have word timingswordsBurns exactly that copy at those times, no speech-to-text

How do you burn captions for a silent opener?

Give each phrase a start and end in seconds. Sume burns exactly that copy; it does not listen to the audio. The call returns a job, so poll GET /v1/jobs/:id/status and then read the result, or use a webhook as described in jobs and results.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: yt-gap-captions-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/opener.mp4",
    "style": "slam",
    "cues": [
      {"text": "Three weeks on the coast", "start": 0.5, "end": 3.0},
      {"text": "No narration, just the sea", "start": 3.2, "end": 6.0}
    ]
  }'

What about overlapping speakers?

Sume's speech-to-text has no documented speaker separation, so do not expect it to untangle a panel talking over itself. If you have a clean script, script_text keeps speech-to-text timings as the timing source and aligns your wording to them. Alignment can fail with script_alignment_mismatch or script_alignment_failed, and the suggested next action is simplify_script_text_or_omit. For hard overlaps, the reliable path is to author cues from a transcript you checked by hand.

If you only need to see what the audio contains first, video inspect can return a transcript at $0.01 per audio minute, and a frames: false call tells you from probe.has_audio whether there is any audio to transcribe.

What does this cost, and what is the catch?

A standalone caption job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog. The catch is that the video is re-rendered with the captions baked in, and the caption source must be a fetchable public HTTPS URL. If you want a toggleable track on YouTube itself, upload a caption file in YouTube Studio; Sume does not produce SRT files from this endpoint. The difference is covered in hard subtitles versus soft subtitles.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume