YouTube automatic captions not available: what to do instead
YouTube lists silence, overlapping speakers and poor audio as reasons auto captions fail. Burn your own with Sume cues or a script_text-aligned caption job.

When YouTube says automatic captions are not available for a video, its own help page gives the causes: audio still processing, an unsupported language, a video over the length limit, poor sound quality or unrecognised speech, a long silence at the start, or several overlapping speakers or languages. You can only fix the first with time. For the rest, a practical route is to burn captions into the file yourself, and Sume's video captions endpoint can do that from your own text and timings.
The YouTube facts here come from Use automatic captioning, read on 2026-10-02. The Sume facts come from the video captions and video inspect docs. Burned-in captions are pixels in the video; they do not populate YouTube's own caption track, so viewers cannot toggle them off, and this post does not claim otherwise.
Why does YouTube skip automatic captions on some videos?
The help page says automatic captions are generated by machine learning, so quality varies, and that mispronunciations, accents, dialects or background noise can make them misrepresent what was said. It also lists the conditions under which captions are not generated at all. Two of them matter most for edited footage: extended silence at the start, which is common in b-roll openers and cinematic intros, and multiple overlapping speakers or simultaneous languages, which is common in panels and reaction videos.
YouTube's advice is to review automatic captions in YouTube Studio under Subtitles and edit the parts that were not transcribed properly. That works when captions exist. When none were produced, you need another source of text and timing.
Which Sume caption input fits each failure?
Sending text fields together is an error: script_text, words, cues and segments are mutually exclusive. A silent clip sent without cues fails as caption_no_speech with next_action: use_overlay_captions, which is the same gap YouTube describes.
| Cause on YouTube's list | Sume caption input | What happens |
|---|---|---|
| Silent clip or long silent opener, no speech at all | cues (or segments) with text, start, end | Burns your authored phrases at your times, no speech-to-text |
| Speech present, you have the script | script_text | Speech-to-text timings stay the source of truth; wording aligns to your script |
| Speech present, no script | Omit all text fields | Speech-to-text wording is burned in |
| You already have word timings | words | Burns exactly that copy at those times, no speech-to-text |
How do you burn captions for a silent opener?
Give each phrase a start and end in seconds. Sume burns exactly that copy; it does not listen to the audio. The call returns a job, so poll GET /v1/jobs/:id/status and then read the result, or use a webhook as described in jobs and results.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: yt-gap-captions-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/opener.mp4",
"style": "slam",
"cues": [
{"text": "Three weeks on the coast", "start": 0.5, "end": 3.0},
{"text": "No narration, just the sea", "start": 3.2, "end": 6.0}
]
}'What about overlapping speakers?
Sume's speech-to-text has no documented speaker separation, so do not expect it to untangle a panel talking over itself. If you have a clean script, script_text keeps speech-to-text timings as the timing source and aligns your wording to them. Alignment can fail with script_alignment_mismatch or script_alignment_failed, and the suggested next action is simplify_script_text_or_omit. For hard overlaps, the reliable path is to author cues from a transcript you checked by hand.
If you only need to see what the audio contains first, video inspect can return a transcript at $0.01 per audio minute, and a frames: false call tells you from probe.has_audio whether there is any audio to transcribe.
What does this cost, and what is the catch?
A standalone caption job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog. The catch is that the video is re-rendered with the captions baked in, and the caption source must be a fetchable public HTTPS URL. If you want a toggleable track on YouTube itself, upload a caption file in YouTube Studio; Sume does not produce SRT files from this endpoint. The difference is covered in hard subtitles versus soft subtitles.
Sources
Related posts
More in Use cases
- YouTube channel trailer: trim three clips, add music, 30 seconds
Build a 30-second channel trailer: trim three best moments at $0.02 each, join them in Timeline 1.0 for $0.10, and add a Lyria music bed at $0.125.
- YouTube inauthentic content rule: a faceless channel checklist
YouTube's monetization policy targets mass-produced, generic content. What a faceless channel should vary per episode, and which Sume steps hold your input.
- YouTube Shorts 3 minutes: Oct 15, 2024 vs Dec 8, 2025 dates
YouTube's help page gives two dates for three-minute Shorts: Oct 15, 2024 for standard channels, Dec 8, 2025 for Artist Channels. How to check a file.
- Zalando AI distortions: a checklist to run before upload
Zalando names three kinds of AI distortion it can reject: wrong article, implausible setting, inconsistent model. A checklist to run on generated images.
Written by Sume