Does the previous agent turn help transcription? AssemblyAI says 4.3%
AssemblyAI reports agent context cut entity error from 16.04% to 15.35%, a 4.3% relative drop. Sume STT has no context field; here is a fair test.

On AssemblyAI's own benchmark page, passing the previous agent turn as context lowers overall entity error from 16.04% to 15.35%. That is 0.69 percentage points, a relative drop of 0.69 / 16.04 = 4.3%. It is a small gain next to the 60.9% and 72.8% the same page reports for word boost on names and technical terms. Sume STT accepts no context or prompt field, so you cannot reproduce the feature; you can only measure whether one matters for your own audio.
Reading the numbers
Two similar-sounding features, two very different effects. Context carries the conversation, so a recogniser can lean toward a likely next phrase. A word list tells the recogniser which rare strings to expect. The first moved entity error by under one point; the second moved name errors by more than half. If you have to choose where to spend effort, the term list has the larger reported effect.
| Feature | Reported effect | Computed |
|---|---|---|
| Agent context (previous turn) | Overall entity error 16.04% to 15.35% | -0.69 points, -4.3% relative |
| Word boost, names | Entity error -60.9% | As reported |
| Word boost, technical terms | Entity error -72.8% | As reported |
| Sume STT | No context or vocabulary field | Post-process in your code |
A fair test without a context field
Run two arms on the same clips. Arm A is raw Sume STT output. Arm B is the same output after a correction step that uses the words of the agent's previous turn: any word in the transcript that is within a small edit distance of a word the agent just said gets replaced by that word. Score both arms with the entity recall scorer from the earlier post. If Arm B does not beat Arm A by a clear margin, context was not your problem.
import difflib, re
def with_context(text, prev_agent_turn, cutoff=0.85):
vocab = {w.lower(): w for w in re.findall(r"[A-Za-z]{4,}", prev_agent_turn)}
def swap(m):
hit = difflib.get_close_matches(m.group(0).lower(), vocab, n=1, cutoff=cutoff)
return vocab[hit[0]] if hit else m.group(0)
return re.sub(r"[A-Za-z]{4,}", swap, text)
print(with_context("my order is for Sumee Studio", "Is that for Sume Studio?"))What to conclude
With a 4.3% relative effect on the vendor's own data, expect little. If your audio is short customer replies to a scripted prompt, the previous turn names the exact words you expect, so the gain could be larger for you. That is exactly why you measure on your own clips. Ten clips cost 10 x $0.01 = 10 cents at one minute each.
Report the result with the clip count and the date, and keep the clips so you can rerun them.
Where this fits in a voice product
Context helps most when the next user turn is predictable: a yes or no, a date, a name the agent just read out. It helps least for open-ended speech. If your product records a customer answering a scripted question and you then transcribe the file, you know the question, so a correction step that uses its words is cheap to build. If your product records free conversation, spend your effort on a term list and on clean audio instead.
In both cases the order of work is the same: measure the baseline on your clips, change one thing, measure again, and keep the change only if the entity recall improves by more than the run-to-run noise you saw when you repeated the baseline.
Sources
Related posts
More in Developers
- Download a 4K Gemini Omni clip from /content: redirect, 409 and retry
GET /v1/videos/{id}/content answers 302 to the file, or 409 job_not_completed while rendering. Python that reads the redirect without leaking your API key.
- Does Dr. split a Sume STT segment? Abbreviations and sentence mode
Segmentation tests each token for a final period, so a token like Dr. or U.S. can end a segment early. Merge on your side with a small abbreviation list.
- Draft with sume/auto, finish on a pinned model: why the look differs
sume/auto never says which model ran, so a draft from it is no preview for a pinned final. Draft and finish on the same catalog id, and where Auto fits.
- 'duration_seconds must match duration': send one field, not two
Video Router accepts duration or duration_seconds as the same value. Send both with different numbers and you get a 400. Send one, or send two equal ones.
Written by Sume