Lesson recording captions: Sume STT, 8 minutes for 28 cents

Caption an 8-minute lesson recording on Sume: transcribe at $0.01 per minute (10-minute cap) then burn captions for $0.20. Chain and limits, read 2026-10-10.

4 min readSume
All posts

An 8-minute lesson can be transcribed and captioned for $0.28 on Sume: speech-to-text at $0.01 per minute ($0.08) and one caption job ($0.20). Speech-to-text accepts audio up to 10 minutes, so an 8-minute lesson fits in one request.

Most tutors already have the recording. The two reasons to run it through Sume are a transcript you can edit and burned-in captions that make the video readable muted. The caption job can also transcribe by itself, so the choice is whether you want to see the words first. Prices come from the Sume pricing code (read 2026-10-10).

What is the endpoint chain?

The chain depends on whether the recording is already a Sume-hosted file.

  • Host the file first. Timeline, trim and caption jobs read media.sume.com URLs. A generated video already is one, and /v1/media-imports takes only TikTok and Instagram links, so a lesson recorded elsewhere needs a Sume-hosted copy first.
  • POST /v1/stt-1.0/transcribe with the audio, timestamps on, to get words and sentences. Read the transcript and fix names.
  • POST /v1/video-captions with words (your corrected list) and your chosen style. Send only one of script_text, words, cues or segments.
  • If the lesson is over 10 minutes, cut it first with POST /v1/video-trim at $0.02 per cut and transcribe the parts.

What does it cost?

Speech-to-text scales per minute and captions are per job. A recording over 10 minutes has to be split before STT, so the 20-minute row assumes two 10-minute parts.

Lesson captioning cost (read 2026-10-10)
LengthSTTCaptionsTotal
4 min$0.04$0.20$0.24
8 min$0.08$0.20$0.28
10 min (cap)$0.10$0.20$0.30
20 min as two parts$0.20$0.40$0.60

What are the limits?

STT works on the audio, not on a video's picture, so a recording with a muted track produces nothing. A clip with audio but no speech returns caption_no_speech.

The corrected word list is the one that burns. If you change a word, keep its timing; moving a word's start by hand changes when it appears. The Video captions page covers how words differ from segments. Also, captions go on the video, not the audio, so audio-only lessons need a still plus Timeline first, as in chapter markers from STT sentence segments.

Check the model page for the language field before captioning non-English speech.

When is this the wrong tool?

For a live lesson with a screen share, a platform's own captions are free. Sume is for the version you publish afterwards: clean, branded captions, in your own file. Pair it with an AI avatar for online course videos when new lessons are made, not recorded.

A review step that pays for itself

Spend two minutes reading the transcript before captioning. Names, product terms and numbers are where speech-to-text most often slips, and one wrong word on screen is more visible than one wrong word in a transcript.

Because the caption job is a flat $0.20, re-burning after a fix costs the same as the first run, so batch your corrections and burn once.

  • Fix names and numbers first.
  • Keep timing intact when you edit a word.
  • Choose one style for the whole course so lessons look the same.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume