What is a partial transcript in streaming speech to text?

A partial is a provisional transcript a streaming model revises as audio arrives. Why subtitles for a finished clip only need final text and word times.

5 min readSume
All posts

A partial transcript is a provisional guess at what a speaker has said so far, sent while the person is still talking. A streaming speech-to-text model keeps revising the guess as more audio arrives, then emits a final transcript once the phrase is settled. Microsoft's announcement for MAI-Transcribe-2-Streaming says the model "produces its first hypotheses (known as 'partials') in just over 100ms" and that it ranks first on Artificial Analysis for both final and partial transcripts (read 2026-10-07).

If you are adding subtitles to a clip that already exists, you almost never need partials. You have the whole recording, so you can wait for the final text and burn it once. Partials matter when something must react before the sentence ends.

Partial versus final in plain terms

Think of a partial as a draft that can change. The word at the end of a partial is the one most likely to be rewritten when the next half-second of audio lands. A final is a draft the model has committed to. Microsoft describes the streaming variant as one that generates partial transcripts as audio arrives, so a voice agent can act mid-sentence, while the batch model processes complete audio after recording ends (read 2026-10-07).

That difference is about when text becomes available, not about what a finished transcript looks like. Both end with text and timings for the same speech.

Partial and final text, as described on Microsoft's streaming page (read 2026-10-07)
PropertyPartialFinal
ArrivesWhile the speaker is mid-phraseAfter the phrase is settled
Can changeYes, as more audio arrivesNo, committed
Use it forVoice agents reacting mid-sentence, live displayStored transcripts, subtitles, search
Cost of waitingNone: it is the early signalDelay until the phrase ends

Who needs partials

Partials are an interaction tool. A voice agent that must start thinking before you finish talking uses them. A live caption overlay that should show words as they are spoken uses them, and accepts that a word on screen may flip.

A subtitle track for a recorded Short, Reel or TikTok has none of those constraints. Nobody is waiting on a half-spoken sentence, and a flickering caption that changes after it appears is exactly what you do not want to ship.

  • Voice agent that interrupts or answers mid-sentence: partials help.
  • Live on-screen captions during an event: partials help, with some word flicker.
  • Subtitles burned into a finished video: use final text and word times.
  • Searchable transcript of an archive: use final text.

What Sume returns for a recorded clip

Sume STT 1.0 is a job, not a stream. You send a public HTTPS audio_url to POST /v1/stt-1.0/transcribe, and the completed result carries text and words[], each with a word, a start and an end in seconds from the audio start. Word timings are always returned. Adding segmentation: {"mode": "sentence"} also returns gapless sentence segments.

Sume webhooks follow the same philosophy: they send terminal job events only, with no progress or partial deliveries, per the webhooks docs. A job.completed event means the final text is ready to read.

Sume lists STT at $0.01 per audio minute, or $0.60 per hour, and the live rate is in GET /v1/catalog. Microsoft's introductory streaming rate is $0.54 per hour through the end of the year (read 2026-10-07). Those two numbers answer a different question than latency does.

A simple rule

Ask when the text is consumed. If a human or a script reads it only after the clip is finished, take final text and burn it with caption cues or a caption job. If something has to act while the audio is still playing, you want a streaming model and its partials, and you should treat any text on screen as provisional until the final arrives.

Most subtitle buyers are in the first group, which is why a 100 millisecond first-hypothesis figure is interesting for agents and not very relevant for a Shorts caption workflow.

Common confusions

Partial does not mean low quality. A partial is early, and its accuracy is judged separately from the final. Microsoft's page reports a ranking for both, which suggests the vendor treats them as two outputs of one model, not one good output and one rough one.

Streaming also does not mean live video. A streaming model reads audio as it arrives; it says nothing about captions appearing on a video player. Putting text on a published video is a render step, and Sume does that as a caption job on a finished file, at $0.20 for a video of up to 60 seconds. Sume documents no streaming transcription endpoint and no real-time captioning, so for a live event you would pair a streaming provider with your own display layer.

Finally, partial results are not free to store. If you log every revision you create many near-duplicate rows. Most systems keep only the final per utterance, plus a timestamp of when it was committed.

What to do next

If you only ever process finished files, skip streaming and read the batch versus streaming decision for a short checklist. If you are exploring both, run a handful of your own clips through each path, compare final text and word times, and decide on measured differences, not leaderboard positions. Sume's job returns words[] with start and end on every call, so the comparison is a few lines of code.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume