Caption files: YouTube UTF-8, LinkedIn text-only SRT, Sume burns in
YouTube takes SRT, SBV and more as UTF-8; LinkedIn wants text-only SRT. Sume burns captions in and exports no SRT, so build the file from segments.

YouTube accepts caption files in several formats including .srt and .sbv, as plain UTF-8, and LinkedIn video ads want captions as an SRT file with text only. Sume does not export an SRT, and its caption model does not take one. It burns text into the picture. For a sidecar file, take the sentence segments from video inspect and write the SRT yourself.
What the platforms ask for
The facts below are from the YouTube caption file page and the LinkedIn video ads page, both read on 2026-10-08.
| Platform | Accepted | Rule |
|---|---|---|
| YouTube | srt, sbv or sub, mpsub, lrc, cap | Plain text, UTF-8; basic types carry no style information |
| LinkedIn video ads | srt | Text only |
What Sume gives you
Video captions burns styled text onto a public HTTPS video for $0.20 per job on videos up to 60 seconds. It does not accept SRT upload. Video inspect returns a transcript when you ask for it, and with segmentation set to sentence each segment has index, text, start, end and duration_seconds. Speech-to-text is $0.01 per audio minute plus Modal compute.
That is enough to build the file. The script below turns segments into SRT text. It uses sample data in the same shape; replace it with the segments you read from a finished inspect job.
def stamp(t):
ms = round(t * 1000)
h, ms = divmod(ms, 3600000)
m, ms = divmod(ms, 60000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
segments = [
{"index": 0, "text": "Welcome back.", "start": 0.0, "end": 1.4},
{"index": 1, "text": "Three specs changed.", "start": 1.4, "end": 3.2},
]
lines = []
for n, seg in enumerate(segments, 1):
lines.append(str(n))
lines.append(stamp(seg["start"]) + " --> " + stamp(seg["end"]))
lines.append(seg["text"])
lines.append("")
open("captions.srt", "w", encoding="utf-8").write("\n".join(lines))Which route for which job
Use the sidecar SRT when the platform shows its own caption control, as YouTube does, because viewers can turn it off and the platform can translate it. Use burned-in captions when the placement has no caption control, or when you want a styled look. The two are not exclusive: you can upload the SRT and also ship a burned-in version for placements that need it.
Name the file after the video so the pair is easy to match later. Keep the file plain. YouTube's page says the basic types hold no style information, and LinkedIn asks for text only, so leave out tags and positioning.
- Sidecar SRT: free to build once you have segments; no Sume caption job.
- Burned-in: $0.20 per job up to 60 seconds, style such as punch or tiktok-green.
- Both: inspect transcript once, then reuse the same segments for SRT and for cues.
Check the encoding
Write the file as UTF-8 with no byte order surprises, which Python's encoding argument above handles. Open it in a text editor and confirm that accented or Hangul text reads correctly before you upload, since a wrong encoding is the most common reason a caption file is refused.
Sources
Related posts
More in Developers
- caption_no_speech: caption a silent clip with cues ($0.20)
Sume's caption job fails with caption_no_speech on silent clips. Send cues with text, start and end in seconds to burn authored text without speech-to-text.
- Cartesia Line SDK hosting ends Dec 1 2026: does batch TTS change?
Cartesia's changelog says Line SDK agent hosting ends Dec 1, 2026 and agent LLM charges began Oct 1. What that means, and what it does not for batch TTS jobs.
- Cartesia sonic-3-latest and sonic-3-preview aliases: fix model strings
Cartesia lists sonic-3-latest and sonic-3-preview as deprecated aliases. What each maps to, which ids Sume's TTS Router accepts, and a safe migration.
- Celery task for a Sume video: task id as the Idempotency-Key
One Celery task submits POST /v1/videos with its own task id as the Idempotency-Key and re-queues itself every 30 seconds until the job ends. No double bills.
Written by Sume