Entity error 14.41%? Score your own call audio with Sume STT in Python
AssemblyAI reports 14.41% entity error on voice-agent audio and 3.44% English WER. Neither is yours. Compute entity recall on 20 of your clips with Sume STT.

AssemblyAI's benchmarks page reports 14.41% overall entity error for voice agents and 3.44% word error rate on English short-form audio, for its realtime model, from runs between 2026-08-27 and 2026-09-22. Those numbers describe the vendor's test sets. Your call recordings have your accents, your noise and your product names, so score them yourself. Entity recall, the share of must-have strings the transcript contains, takes about 20 lines against Sume STT.
Entity error versus word error
Word error rate counts every wrong word equally, so a missed 'the' costs the same as a missed order number. Entity error counts only the strings that carry meaning, such as names, numbers, codes and addresses, which is why it is a better score for a voice agent. The page reports both kinds of figure, in several settings: 14.41% overall entity error, 6.13% WER in heavy noise, and 6.64% entity error on Indian-accented speech.
| Setting | Metric | Reported figure |
|---|---|---|
| Voice agents, overall | Entity error | 14.41% |
| Heavy noise | WER | 6.13% |
| Indian-accented speech | Entity error | 6.64% |
| English short-form | WER | 3.44% |
| Code-switching, five language pairs | WER | 7.20% |
A scorer for your own clips
Make a CSV of clip URLs and the strings that must appear in each. Transcribe each clip with Sume STT, flatten every string in the JSON result, and test whether each must-have string appears after removing case, spaces and punctuation. Entity recall is hits divided by total.
import os, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
def run(path, body):
d = requests.post(B + path, json=body, headers=H, timeout=60).json()["data"]
while not d.get("terminal"):
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
return requests.get(d["result_url"], headers=H, timeout=60).json()
def flat(x):
if isinstance(x, str):
return x
vals = x.values() if isinstance(x, dict) else x if isinstance(x, list) else []
return " ".join(flat(v) for v in vals)
norm = lambda s: "".join(c for c in s.lower() if c.isalnum())
clips = [("https://example.com/a.mp3", ["4F7Q92", "Elm Street"])]
hit = total = 0
for url, must in clips:
blob = norm(flat(run("/v1/stt-1.0/transcribe", {"audio_url": url, "duration_seconds": 60})))
for m in must:
total += 1
hit += norm(m) in blob
print(f"entity recall {hit}/{total}")Reading your own number
Twenty one-minute clips cost 20 x $0.01 = 20 cents to transcribe. Report the recall, then list the misses by type: digits, names or addresses. If digits dominate, a spoken-number normaliser is the fix; if names dominate, use a term map. Compare the result with the vendor's figure only as a loose guide, since the test sets differ.
Add noisy clips on purpose, since the page reports its noise result separately for the same reason.
Building the clip set
Pick 20 clips that look like production: some quiet, some noisy, some with strong accents, some with fast speech. Under 10 minutes each, because Sume STT takes up to 600 seconds per job, and with a public HTTPS URL. Write the must-have strings by listening, not by reading a transcript, so the answer key does not inherit the recogniser's mistakes.
Keep the set fixed. When you try another engine or a new model version next quarter, run the same clips and compare entity recall on identical input, which is the only comparison that means anything.
What the scorer does not see
Substring matching is forgiving: it will count '4F7Q92' as found inside a longer string that merely contains it, and it cannot tell a correct reading from a lucky one. For short codes, check that neighbouring words are also right, or compare against the word timings for the clip. For a first pass, though, a simple recall number is enough to tell you whether to worry.
Sources
Related posts
More in Developers
- Assign a Sume video model to each old prompt by clip length
A short Python planner that reads a CSV of old prompt lengths, sends 10 seconds or less to Omni, up to 30 seconds to Wan 3.0, and splits longer clips.
- Audio detach errors: unsupported_media_type, source_not_found
Each Sume audio detach refusal code and its one-line fix: off-host URL, other workspace, not a video, empty range, no audio track, source too long.
- automation_generation_spend_cap_exceeded: what a 402 cap error ends
The per-run cap rejects only the one generation that would cross it. current, requested and cap come back in micros so you can size the next request.
- Avatar batch: which limit hits first, writes per minute or the queue?
On Pro, 300 writes/min is far above the 24 accepted jobs (4 running, 20 queued). Queue capacity limits an avatar batch first, so submit in waves of that size.
Written by Sume