Subtitles for a recorded video: 100 ms partials or final text?

MAI-Transcribe-2-Streaming sells 100 ms partials for live subtitling. For a recorded video, final text and word timings matter more. Choosing with Sume STT.

5 min readSume
All posts

For a video you have already recorded, you do not need partial transcripts. You need the final text, correct word timings and a render. Microsoft's MAI-Transcribe-2-Streaming returns first partials just over 100 ms after audio arrives and revises them as context rolls in, which is valuable when a person is speaking right now. On a finished file, wait for the whole result: Sume STT returns text and words[] with start and end seconds once the job completes.

Live and recorded are different jobs

The facts in this table are from the Microsoft announcement, the OpenAI speech guide and the Sume API description, read on 2026-10-04.

Streaming and file transcription compared (read 2026-10-04)
QuestionLive captionsRecorded video
What matters mostTime to first wordsAccuracy of the final text
Typical serviceStreaming recogniser (MAI-Transcribe-2-Streaming, 60 languages)File job (Sume STT 1.0, or OpenAI file transcription)
RevisionsPartials change as context arrivesOne final transcript
Word timingsNeeded to place words as spokenNeeded to burn or export captions
Price shapePer hour of audio ($0.54 per hour intro rate)Per audio minute (Sume $0.01)

What a file job gives you

Sume's POST /v1/stt-1.0/transcribe returns text, language fields where available, and words[]. Add segmentation: { mode: "sentence" } and you also get sentence segments, which are time ranges over your audio. There is no flag for timestamps because they are always returned. A job is async by default; the sync mode waits at most 30 seconds, and the jobs and results page shows the polling side. OpenAI's guide adds a useful detail for subtitle work: only whisper-1 supports word and segment timestamps there, while gpt-transcribe is the recommended model for plain transcription.

A decision rule

Ask whether a human watches words appear while the speaker is still talking. If yes, you need a streaming service; Microsoft's claim that words appear two times faster than its closest competitor is a vendor claim, so test it on your audio. If no, a file job is simpler, because you hand over one URL and read one result.

Many teams need both: live captions during the event, then a clean burned-in version for the replay. For the second, the video captions job transcribes the clip and renders styled captions in one call at $0.20 per job for clips up to 60 seconds. Use script_text if the final wording must match a known script.

What not to expect

A file job will not give you the first words of a speaker in 100 ms, and a streaming service will not render a styled caption onto your video. Neither replaces a human proofread for names and numbers. Budget a read-through of the final transcript before you burn. If a typo slips through, a restyle with source_caption_id and corrected words re-burns it, but that is another $0.20 render.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume