What is a partial transcript in streaming speech to text?
A partial is a provisional transcript a streaming model revises as audio arrives. Why subtitles for a finished clip only need final text and word times.

A partial transcript is a provisional guess at what a speaker has said so far, sent while the person is still talking. A streaming speech-to-text model keeps revising the guess as more audio arrives, then emits a final transcript once the phrase is settled. Microsoft's announcement for MAI-Transcribe-2-Streaming says the model "produces its first hypotheses (known as 'partials') in just over 100ms" and that it ranks first on Artificial Analysis for both final and partial transcripts (read 2026-10-07).
If you are adding subtitles to a clip that already exists, you almost never need partials. You have the whole recording, so you can wait for the final text and burn it once. Partials matter when something must react before the sentence ends.
Partial versus final in plain terms
Think of a partial as a draft that can change. The word at the end of a partial is the one most likely to be rewritten when the next half-second of audio lands. A final is a draft the model has committed to. Microsoft describes the streaming variant as one that generates partial transcripts as audio arrives, so a voice agent can act mid-sentence, while the batch model processes complete audio after recording ends (read 2026-10-07).
That difference is about when text becomes available, not about what a finished transcript looks like. Both end with text and timings for the same speech.
| Property | Partial | Final |
|---|---|---|
| Arrives | While the speaker is mid-phrase | After the phrase is settled |
| Can change | Yes, as more audio arrives | No, committed |
| Use it for | Voice agents reacting mid-sentence, live display | Stored transcripts, subtitles, search |
| Cost of waiting | None: it is the early signal | Delay until the phrase ends |
Who needs partials
Partials are an interaction tool. A voice agent that must start thinking before you finish talking uses them. A live caption overlay that should show words as they are spoken uses them, and accepts that a word on screen may flip.
A subtitle track for a recorded Short, Reel or TikTok has none of those constraints. Nobody is waiting on a half-spoken sentence, and a flickering caption that changes after it appears is exactly what you do not want to ship.
- Voice agent that interrupts or answers mid-sentence: partials help.
- Live on-screen captions during an event: partials help, with some word flicker.
- Subtitles burned into a finished video: use final text and word times.
- Searchable transcript of an archive: use final text.
What Sume returns for a recorded clip
Sume STT 1.0 is a job, not a stream. You send a public HTTPS audio_url to POST /v1/stt-1.0/transcribe, and the completed result carries text and words[], each with a word, a start and an end in seconds from the audio start. Word timings are always returned. Adding segmentation: {"mode": "sentence"} also returns gapless sentence segments.
Sume webhooks follow the same philosophy: they send terminal job events only, with no progress or partial deliveries, per the webhooks docs. A job.completed event means the final text is ready to read.
Sume lists STT at $0.01 per audio minute, or $0.60 per hour, and the live rate is in GET /v1/catalog. Microsoft's introductory streaming rate is $0.54 per hour through the end of the year (read 2026-10-07). Those two numbers answer a different question than latency does.
A simple rule
Ask when the text is consumed. If a human or a script reads it only after the clip is finished, take final text and burn it with caption cues or a caption job. If something has to act while the audio is still playing, you want a streaming model and its partials, and you should treat any text on screen as provisional until the final arrives.
Most subtitle buyers are in the first group, which is why a 100 millisecond first-hypothesis figure is interesting for agents and not very relevant for a Shorts caption workflow.
Common confusions
Partial does not mean low quality. A partial is early, and its accuracy is judged separately from the final. Microsoft's page reports a ranking for both, which suggests the vendor treats them as two outputs of one model, not one good output and one rough one.
Streaming also does not mean live video. A streaming model reads audio as it arrives; it says nothing about captions appearing on a video player. Putting text on a published video is a render step, and Sume does that as a caption job on a finished file, at $0.20 for a video of up to 60 seconds. Sume documents no streaming transcription endpoint and no real-time captioning, so for a live event you would pair a streaming provider with your own display layer.
Finally, partial results are not free to store. If you log every revision you create many near-duplicate rows. Most systems keep only the final per utterance, plus a timestamp of when it was committed.
What to do next
If you only ever process finished files, skip streaming and read the batch versus streaming decision for a short checklist. If you are exploring both, run a handful of your own clips through each path, compare final text and word times, and decide on measured differences, not leaderboard positions. Sume's job returns words[] with start and end on every call, so the comparison is a few lines of code.
Sources
Related posts
More in Developers
- What to show a viewer while an avatar video job is queued
Avatar jobs on Sume are async: queued, processing, then completed, failed or canceled. A status-to-UI map for waiting screens, with polling rules.
- When is async TTS the right choice? Sync wait, poll or webhook
Async TTS is right for voiceovers, batches and anything a person is not watching a spinner for. Sume's sync wait stops at 30 seconds; Flash claims 45 ms.
- Which Sume API calls are safe to retry blindly, and which need a key?
Reads, cancels and redelivers retry safely; paid submits retry only under the same Idempotency-Key. A call-by-call table, plus the codes that mean wait or stop.
- Which Sume job and run endings send a webhook, and which stay silent?
Jobs send job.completed, job.failed and job.canceled. Canceled or skipped runs send nothing. A matrix of terminal states and what your receiver can expect.
Written by Sume