Descript embedded SRT toggle vs Sume burned-in captions
Descript added a toggle for the embedded SRT track on export. Sume's caption job burns text into pixels and does not accept SRT uploads. When each fits.

Descript's changelog for September 17 says exports have a toggle for the embedded SRT track, which was previously always included. An embedded track is a caption stream the viewer's player can switch off. Sume's POST /v1/video-captions does the other thing: it burns captions into the video frames, and its docs state that SRT uploads are not supported. Pick on whether the platform should be able to hide, restyle or translate the captions.
What each product does
The Descript line is from its public changelog entry; the Sume lines are from the Sume video-captions docs.
| Question | Descript | Sume video captions |
|---|---|---|
| Caption form | Embedded SRT track, now optional on export | Burned-in, rendered styles |
| SRT input | Not covered here | Not supported |
| Style set | Redesigned captions with new presets and animations | slam, punch, tiktok-green, korean-ad, black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis |
| Price | See Descript | $0.20 per job for videos up to 60 seconds |
Burned-in with your own text
If you already have timed text, pass it as cues (phrase-level text, start, end) or words. That skips speech recognition and burns your lines as given. Only one of script_text, words, cues or segments can be sent.
import os, requests
r = requests.post(
"https://api.sume.com/v1/video-captions",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "cap-cues-001"},
json={"video_url": "https://example.com/clip.mp4",
"style": "highlight",
"cues": [{"text": "Free shipping today", "start": 0.0, "end": 2.2}]},
)
print(r.status_code, r.json())Choosing
Use an embedded track when platforms or accessibility checks want a separate, toggleable stream. Use burned-in when the platform strips tracks (many short-form feeds) or the caption style is part of the look. If you need both, produce the SRT in your editor and the burned-in version through Sume.
Worked example
The question is usually where the video will be posted.
- A platform that reads sidecar or embedded tracks: an embedded SRT can help, and Descript now lets you leave it off.
- A feed that strips tracks: only burned-in captions will be seen.
- A mix of both: export the track from your editor and run the burned-in version through Sume.
Checklist before you commit
Sume's caption styles all render into the video. If you need to switch captions off for viewers, burned-in is the wrong choice; if you need them to always show, it is the right one.
- Check the host's caption-track support.
- Check Korean or other non-Latin copy against the style you choose; some styles error on mismatched text.
- Keep the source text file for future edits.
Sources
Related posts
More in Comparisons
- Descript music at a set length vs Sume music length in the prompt
Descript's Sept 17 update generates music and effects at specified lengths. Sume Music has no duration field; you ask for length in the prompt. How to do that.
- Descript per-second smoothing credits vs flat API job prices
Descript now bills smoothing by the second. Here is how that compares with Sume media jobs, which charge a flat amount per job or per output minute.
- Dreamina Long Video Mode: 3 minutes vs Sume's 30 s Seedance
Dreamina says its Seedance 2.5 Long Video Mode makes videos up to three minutes. Sume's seedance-2.5 makes 4-30 s per request, so long films are joined clips.
- ElevenLabs Avatars wants 3 to 5 reference images: what Sume needs
ElevenLabs recommends 3-5 reference images from different angles for Avatars. Sume Avatar 1.0 builds an avatar handle from one photo, a prompt, or simple props.
Written by Sume