Subtitles for a recorded video: 100 ms partials or final text?
MAI-Transcribe-2-Streaming sells 100 ms partials for live subtitling. For a recorded video, final text and word timings matter more. Choosing with Sume STT.

For a video you have already recorded, you do not need partial transcripts. You need the final text, correct word timings and a render. Microsoft's MAI-Transcribe-2-Streaming returns first partials just over 100 ms after audio arrives and revises them as context rolls in, which is valuable when a person is speaking right now. On a finished file, wait for the whole result: Sume STT returns text and words[] with start and end seconds once the job completes.
Live and recorded are different jobs
The facts in this table are from the Microsoft announcement, the OpenAI speech guide and the Sume API description, read on 2026-10-04.
| Question | Live captions | Recorded video |
|---|---|---|
| What matters most | Time to first words | Accuracy of the final text |
| Typical service | Streaming recogniser (MAI-Transcribe-2-Streaming, 60 languages) | File job (Sume STT 1.0, or OpenAI file transcription) |
| Revisions | Partials change as context arrives | One final transcript |
| Word timings | Needed to place words as spoken | Needed to burn or export captions |
| Price shape | Per hour of audio ($0.54 per hour intro rate) | Per audio minute (Sume $0.01) |
What a file job gives you
Sume's POST /v1/stt-1.0/transcribe returns text, language fields where available, and words[]. Add segmentation: { mode: "sentence" } and you also get sentence segments, which are time ranges over your audio. There is no flag for timestamps because they are always returned. A job is async by default; the sync mode waits at most 30 seconds, and the jobs and results page shows the polling side. OpenAI's guide adds a useful detail for subtitle work: only whisper-1 supports word and segment timestamps there, while gpt-transcribe is the recommended model for plain transcription.
A decision rule
Ask whether a human watches words appear while the speaker is still talking. If yes, you need a streaming service; Microsoft's claim that words appear two times faster than its closest competitor is a vendor claim, so test it on your audio. If no, a file job is simpler, because you hand over one URL and read one result.
Many teams need both: live captions during the event, then a clean burned-in version for the replay. For the second, the video captions job transcribes the clip and renders styled captions in one call at $0.20 per job for clips up to 60 seconds. Use script_text if the final wording must match a known script.
What not to expect
A file job will not give you the first words of a speaker in 100 ms, and a streaming service will not render a styled caption onto your video. Neither replaces a human proofread for names and numbers. Budget a read-through of the final transcript before you burn. If a typo slips through, a restyle with source_caption_id and corrected words re-burns it, but that is another $0.20 render.
Sources
Related posts
More in Comparisons
- Suno Premier's 30-minute upload versus Sume's audio limits
Suno Premier allows audio uploads up to 30 minutes. Sume's Timeline audio joins up to 1,800 s (30 min), while audio detach takes 1,800 s in and 900 s out.
- Suno Speech beta: voice and music in one pass, or separate tracks?
Suno's Speech beta makes voice and music in one track. Its blog lists wandering accents and long pauses. When to prefer separate TTS, music and a timeline mix.
- Suno terms: commercial use needs an approved Pro or Premier download
Suno terms limit free and basic outputs to personal use; Pro and Premier may use them commercially via an approved download. What that means for a video ad.
- Suno v6 and label partners: what it means for an ad soundtrack
Suno's v6 post names WMG, BMG and Believe and upload screening. What a rights-minded team should check before using any AI music in a paid ad.
Written by Sume