Sonix SRT and VTT export vs Sume burned-in captions

Need a subtitle file? Sonix exports SRT and VTT; Sume video captions return a captioned video, not a sidecar file. Compare outputs, languages and $0.20 jobs.

5 min readSume
All posts

If you need a subtitle file you can upload to a platform, Sonix is built for it and Sume video captions are not. Sonix's features page (read 2026-10-10) lists SRT and VTT export plus burned-in subtitles, while Sume's caption job burns styled words onto your video and returns a captioned video_url.

The Sume docs say the resource returns the status, the style and the captioned video, and that raw transcripts are not part of the public contract. That is a clear boundary, so the choice is about which output your platform needs.

Output and pricing at a glance

Sonix sells transcription by the hour with monthly plans; Sume sells a fixed caption job. Sonix's page calls out 54+ languages and "99% accuracy"; Sume's caption docs describe a language hint for speech-to-text and do not state a language count.

Sonix from its pricing and features pages (read 2026-10-10); Sume from the video captions docs and API description
ItemSonix (read 2026-10-10)Sume video captions
OutputSRT, VTT and burned-in subtitlesA captioned MP4 (video_url)
Speaker labelsBuilt into transcriptsNot part of the public result
Languages54+language hint; auto-detect if omitted
PricePay as you go $10 per hour; Core $25 per month for 5 hours$0.20 per job for videos up to 60 seconds
Free trial30 minutes, no cardNone stated
Text without speechNot statedcues or segments with text, start and end
InputRESTful API with webhooksPublic HTTPS video_url

What a Sume caption job does well

The strength is the look. Standalone captions take a style such as slam (the Latin default), punch or tiktok-green, plus Hangul styles for Korean speech. A design object changes colors, typography, placement, phrasing and motion for one request, and out-of-range values fail with a 400 at request time rather than producing a bad render you pay for.

If you hold the exact wording, send script_text to lock the burned-in text to your script while speech-to-text timings stay the timing source. If alignment cannot map the script onto the speech, the job fails with script_alignment_mismatch or script_alignment_failed. For a clip with no speech, pass cues instead; otherwise the job fails with caption_no_speech.

The gap: no sidecar file

Platforms that accept a separate subtitle track, so viewers can toggle it, want SRT or VTT. Sume's caption docs do not describe producing one, and say the API does not support SRT uploads either. A burned-in caption is permanent and cannot be turned off by the viewer or translated by the platform.

One path exists on the Sume side: speech-to-text always returns word timings, with optional sentence segments, so you could build a subtitle file from that output in your own code. I am not claiming Sume ships that as a feature, only that the timing data is in the STT result.

Choosing

Use a short list.

  • Viewers should toggle or translate captions: you need SRT or VTT, so Sonix.
  • Short vertical video where captions are part of the design: Sume captions at $0.20 per clip.
  • Both: transcribe and export a file elsewhere, and burn the styled version with Sume.
  • Clips longer than 60 seconds: the $0.20 figure covers up to 60 seconds under the current estimate, so check the live rate before batching long videos.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume