Does the previous agent turn help transcription? AssemblyAI says 4.3%

AssemblyAI reports agent context cut entity error from 16.04% to 15.35%, a 4.3% relative drop. Sume STT has no context field; here is a fair test.

5 min readSume
All posts

On AssemblyAI's own benchmark page, passing the previous agent turn as context lowers overall entity error from 16.04% to 15.35%. That is 0.69 percentage points, a relative drop of 0.69 / 16.04 = 4.3%. It is a small gain next to the 60.9% and 72.8% the same page reports for word boost on names and technical terms. Sume STT accepts no context or prompt field, so you cannot reproduce the feature; you can only measure whether one matters for your own audio.

Reading the numbers

Two similar-sounding features, two very different effects. Context carries the conversation, so a recogniser can lean toward a likely next phrase. A word list tells the recogniser which rare strings to expect. The first moved entity error by under one point; the second moved name errors by more than half. If you have to choose where to spend effort, the term list has the larger reported effect.

AssemblyAI features and reported effect, benchmarks page, read 2026-10-05
FeatureReported effectComputed
Agent context (previous turn)Overall entity error 16.04% to 15.35%-0.69 points, -4.3% relative
Word boost, namesEntity error -60.9%As reported
Word boost, technical termsEntity error -72.8%As reported
Sume STTNo context or vocabulary fieldPost-process in your code

A fair test without a context field

Run two arms on the same clips. Arm A is raw Sume STT output. Arm B is the same output after a correction step that uses the words of the agent's previous turn: any word in the transcript that is within a small edit distance of a word the agent just said gets replaced by that word. Score both arms with the entity recall scorer from the earlier post. If Arm B does not beat Arm A by a clear margin, context was not your problem.

import difflib, re

def with_context(text, prev_agent_turn, cutoff=0.85):
    vocab = {w.lower(): w for w in re.findall(r"[A-Za-z]{4,}", prev_agent_turn)}
    def swap(m):
        hit = difflib.get_close_matches(m.group(0).lower(), vocab, n=1, cutoff=cutoff)
        return vocab[hit[0]] if hit else m.group(0)
    return re.sub(r"[A-Za-z]{4,}", swap, text)

print(with_context("my order is for Sumee Studio", "Is that for Sume Studio?"))

What to conclude

With a 4.3% relative effect on the vendor's own data, expect little. If your audio is short customer replies to a scripted prompt, the previous turn names the exact words you expect, so the gain could be larger for you. That is exactly why you measure on your own clips. Ten clips cost 10 x $0.01 = 10 cents at one minute each.

Report the result with the clip count and the date, and keep the clips so you can rerun them.

Where this fits in a voice product

Context helps most when the next user turn is predictable: a yes or no, a date, a name the agent just read out. It helps least for open-ended speech. If your product records a customer answering a scripted question and you then transcribe the file, you know the question, so a correction step that uses its words is cheap to build. If your product records free conversation, spend your effort on a term list and on clean audio instead.

In both cases the order of work is the same: measure the baseline on your clips, change one thing, measure again, and keep the change only if the entity recall improves by more than the run-to-run noise you saw when you repeated the baseline.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume