YouTube captions.insert: 100 MB, 400 quota units, and an SRT build
YouTube captions.insert costs 400 quota units and takes a 100 MB file. Sume returns words and segments, not SRT, so here is the 20-line conversion to upload.

YouTube's captions.insert method uploads a caption track for a video. The reference gives a 100 MB file limit and a quota cost of 400 units per call, and requires snippet.videoId, snippet.language and snippet.name; snippet.isDraft is optional. Sume's video inspect returns a transcript with timed segments, but not a subtitle file, so you build the SRT yourself. That is about twenty lines and avoids burning text into the picture.
Facts from the method page
| Item | Value |
|---|---|
| Maximum file size | 100 MB |
| Quota cost | 400 units per call |
| Required fields | snippet.videoId, snippet.language, snippet.name |
| Optional field | snippet.isDraft (boolean) |
| Accepted types listed | text/xml, application/octet-stream, */* |
Step 1: get timed segments from Sume
POST /v1/video-inspect with transcribe: true and segmentation: {"mode": "sentence"} returns transcript.segments[], each with index, text, start and end in seconds. Transcription is billed at $0.01 per audio minute, and a silent clip fails with inspect_source_has_no_audio; check probe.has_audio first.
Step 2: write the SRT
def stamp(sec):
ms = round(sec * 1000)
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
def to_srt(segments):
blocks = []
for i, seg in enumerate(segments, 1):
blocks.append(f"{i}\n{stamp(seg['start'])} --> "
f"{stamp(seg['end'])}\n{seg['text'].strip()}\n")
return "\n".join(blocks)
demo = [{"text": "Hello there.", "start": 0.0, "end": 1.4}]
print(to_srt(demo))
Why a track instead of burned-in text
A caption track can be turned off, translated and searched; burned-in text cannot. Sume's video captions endpoint only burns text into the video ($0.20 per clip up to 60 seconds), and the docs mention no SRT or VTT output (and say SRT uploads are unsupported). Use burned-in captions for platforms that cannot take a track, and the SRT route for YouTube.
Putting it together
The flow is: import the clip, probe it, run inspect with transcribe: true, convert transcript.segments with to_srt, then call captions.insert with the video id, a language code and a name. Keep isDraft true on the first run so the track is not live until a person has read it. Segment boundaries come from sentence splitting, so long sentences become long subtitle blocks; if that reads badly on screen, split the text yourself at punctuation before writing the file, and keep each block short enough to read at a glance.
Limits
At 400 units per call, a default quota will not stretch far; check your project quota before a batch. We did not verify YouTube's accepted subtitle formats beyond the MIME list on the page, so confirm SRT with a draft track (isDraft). Transcription quality depends on the audio, and automatic text needs a human pass for names and brands.
Sources
Related posts
More in Developers
- YouTube videos.insert: 256 GB limit and containsSyntheticMedia
The YouTube Data API videos.insert method takes files up to 256 GB and a status.containsSyntheticMedia flag. A request body for an AI-made upload, with checks.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
Written by Sume