YouTube captions.insert: 100 MB, 400 quota units, and an SRT build

YouTube captions.insert costs 400 quota units and takes a 100 MB file. Sume returns words and segments, not SRT, so here is the 20-line conversion to upload.

4 min readSume
All posts

YouTube's captions.insert method uploads a caption track for a video. The reference gives a 100 MB file limit and a quota cost of 400 units per call, and requires snippet.videoId, snippet.language and snippet.name; snippet.isDraft is optional. Sume's video inspect returns a transcript with timed segments, but not a subtitle file, so you build the SRT yourself. That is about twenty lines and avoids burning text into the picture.

Facts from the method page

YouTube captions.insert reference, read 2026-10-02
ItemValue
Maximum file size100 MB
Quota cost400 units per call
Required fieldssnippet.videoId, snippet.language, snippet.name
Optional fieldsnippet.isDraft (boolean)
Accepted types listedtext/xml, application/octet-stream, */*

Step 1: get timed segments from Sume

POST /v1/video-inspect with transcribe: true and segmentation: {"mode": "sentence"} returns transcript.segments[], each with index, text, start and end in seconds. Transcription is billed at $0.01 per audio minute, and a silent clip fails with inspect_source_has_no_audio; check probe.has_audio first.

Step 2: write the SRT

def stamp(sec):
    ms = round(sec * 1000)
    h, ms = divmod(ms, 3_600_000)
    m, ms = divmod(ms, 60_000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

def to_srt(segments):
    blocks = []
    for i, seg in enumerate(segments, 1):
        blocks.append(f"{i}\n{stamp(seg['start'])} --> "
                      f"{stamp(seg['end'])}\n{seg['text'].strip()}\n")
    return "\n".join(blocks)

demo = [{"text": "Hello there.", "start": 0.0, "end": 1.4}]
print(to_srt(demo))

Why a track instead of burned-in text

A caption track can be turned off, translated and searched; burned-in text cannot. Sume's video captions endpoint only burns text into the video ($0.20 per clip up to 60 seconds), and the docs mention no SRT or VTT output (and say SRT uploads are unsupported). Use burned-in captions for platforms that cannot take a track, and the SRT route for YouTube.

Putting it together

The flow is: import the clip, probe it, run inspect with transcribe: true, convert transcript.segments with to_srt, then call captions.insert with the video id, a language code and a name. Keep isDraft true on the first run so the track is not live until a person has read it. Segment boundaries come from sentence splitting, so long sentences become long subtitle blocks; if that reads badly on screen, split the text yourself at punctuation before writing the file, and keep each block short enough to read at a glance.

Limits

At 400 units per call, a default quota will not stretch far; check your project quota before a batch. We did not verify YouTube's accepted subtitle formats beyond the MIME list on the page, so confirm SRT with a draft track (isDraft). Transcription quality depends on the audio, and automatic text needs a human pass for names and brands.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume