Sonix SRT and VTT export vs Sume burned-in captions
Need a subtitle file? Sonix exports SRT and VTT; Sume video captions return a captioned video, not a sidecar file. Compare outputs, languages and $0.20 jobs.

If you need a subtitle file you can upload to a platform, Sonix is built for it and Sume video captions are not. Sonix's features page (read 2026-10-10) lists SRT and VTT export plus burned-in subtitles, while Sume's caption job burns styled words onto your video and returns a captioned video_url.
The Sume docs say the resource returns the status, the style and the captioned video, and that raw transcripts are not part of the public contract. That is a clear boundary, so the choice is about which output your platform needs.
Output and pricing at a glance
Sonix sells transcription by the hour with monthly plans; Sume sells a fixed caption job. Sonix's page calls out 54+ languages and "99% accuracy"; Sume's caption docs describe a language hint for speech-to-text and do not state a language count.
| Item | Sonix (read 2026-10-10) | Sume video captions |
|---|---|---|
| Output | SRT, VTT and burned-in subtitles | A captioned MP4 (video_url) |
| Speaker labels | Built into transcripts | Not part of the public result |
| Languages | 54+ | language hint; auto-detect if omitted |
| Price | Pay as you go $10 per hour; Core $25 per month for 5 hours | $0.20 per job for videos up to 60 seconds |
| Free trial | 30 minutes, no card | None stated |
| Text without speech | Not stated | cues or segments with text, start and end |
| Input | RESTful API with webhooks | Public HTTPS video_url |
What a Sume caption job does well
The strength is the look. Standalone captions take a style such as slam (the Latin default), punch or tiktok-green, plus Hangul styles for Korean speech. A design object changes colors, typography, placement, phrasing and motion for one request, and out-of-range values fail with a 400 at request time rather than producing a bad render you pay for.
If you hold the exact wording, send script_text to lock the burned-in text to your script while speech-to-text timings stay the timing source. If alignment cannot map the script onto the speech, the job fails with script_alignment_mismatch or script_alignment_failed. For a clip with no speech, pass cues instead; otherwise the job fails with caption_no_speech.
The gap: no sidecar file
Platforms that accept a separate subtitle track, so viewers can toggle it, want SRT or VTT. Sume's caption docs do not describe producing one, and say the API does not support SRT uploads either. A burned-in caption is permanent and cannot be turned off by the viewer or translated by the platform.
One path exists on the Sume side: speech-to-text always returns word timings, with optional sentence segments, so you could build a subtitle file from that output in your own code. I am not claiming Sume ships that as a feature, only that the timing data is in the STT result.
Choosing
Use a short list.
- Viewers should toggle or translate captions: you need SRT or VTT, so Sonix.
- Short vertical video where captions are part of the design: Sume captions at $0.20 per clip.
- Both: transcribe and export a file elsewhere, and burn the styled version with Sume.
- Clips longer than 60 seconds: the $0.20 figure covers up to 60 seconds under the current estimate, so check the live rate before batching long videos.
Sources
Related posts
More in Comparisons
- Sony Woosh sound effects model: does Sume have SFX or video-to-audio?
Sony AI's Woosh makes sound effects and video-to-audio. Sume has no SFX route; it offers video models with native audio, plus music stings. What to use instead.
- Spreadsheet rows to videos: Creatomate or Sume bulk runs?
Creatomate ties a template to a spreadsheet; Sume bulk runs queue up to 100 Format runs at concurrency 1 to 16. How rows map, and how many queues 250 rows take.
- Suno Italy probe: four terms clauses to check before ad music
Italy's AGCM opened a probe into Suno's terms on Oct 6, 2026. The four flagged clauses, and what to check in any AI music tool before putting a track in an ad.
- Suno Speech makes voice and music in one pass. What does Sume do?
Suno's Speech beta (Oct 1, 2026) makes voice and music as one track. Sume has no such model; it layers TTS, Music and a Timeline soundtrack. Trade-offs inside.
Written by Sume