Captions app SRT export vs Sume burned-in caption cues
The Captions app can export an SRT file for external caption tracks. Sume burns captions into the MP4 from authored cues with text, start and end.

The Captions app can export a video together with an SRT subtitle file. Sume's caption output is different: text is burned into the video itself, and the docs describe no sidecar file. If a platform needs a separate track, build it from your own cue list.
Captions facts are from its release notes; Sume facts from Video captions and Generate avatar video, read 2026-10-01.
What does the SRT export do?
The release notes say videos can be exported with an SRT subtitle file, useful for platforms that support external caption tracks. The track stays separate from the picture, so a platform can restyle or switch it off.
What does Sume produce instead?
Standalone captions take a public HTTPS video_url and return a captioned video. Inline captions on an avatar video burn styles into the clean final MP4. Neither page describes SRT or VTT output; the caption page lists SRT uploads as unsupported and says to pass phrase-level text as cues instead.
| Question | Captions app | Sume |
|---|---|---|
| Subtitle file | SRT export | Not described in the docs |
| Burned-in text | Caption styles in the app | style on caption job or inline captions |
| Authored text | Not covered | cues with text, start, end |
Can I author the caption text myself?
Yes. Pass cues (or segments) with text, start and end and Sume burns that overlay copy without speech-to-text. That same list is a good source for writing an .srt yourself, since you already hold the timings.
Is there an extra charge for inline captions?
Inline captions do not create a separate billed video-caption job. Standalone captioning of an existing clip is its own job; see our burn captions onto video guide.
Sources
Related posts
More in Use cases
- Captions AI avatar looks vs one Sume avatar handle per look
Captions added Avatar Looks to save new looks per avatar. On Sume you create one avatar handle per look from a prompt, a profile or a reference image.
- Multilingual video captions: language is a hint, not a font
On Sume video captions, `language` only tells speech-to-text what to expect. Look and font come from `style`, `design` and `font`, never from the language.
- Chrome Web Store 440x280 promo tile: generate at 1760x1120
The Chrome Web Store small promo tile is 440x280, below GPT Image 2.5's pixel floor. Request 1760x1120 (exact 4x, same 11:7 ratio) on Sume and downscale.
- Chrome Web Store screenshots at 1280x800: request it directly
Chrome Web Store screenshots are 1280x800 (or 640x400), JPG or PNG, square corners, full bleed. 1280x800 passes GPT Image 2.5 rules, so request it on Sume.
Written by Sume