AssemblyAI word boost cuts name errors 61%: a term map for Sume STT

AssemblyAI reports word boost cut entity errors 60.9% on names and 72.8% on technical terms. Sume STT has no boost field; here is a post-correction map.

5 min readSume
All posts

AssemblyAI's benchmarks page says its word boost feature reduces entity error on names by 60.9% and on technical terms by 72.8%, in tests run from 2026-08-27 to 2026-09-22. Sume STT has no vocabulary or boost field: the request takes audio_url, an optional language_code, duration_seconds, segmentation and delivery fields, and nothing else. What you can do is correct the output afterwards, using a term map of the names and product words you already know.

What the benchmark says, and the caveat

The page reports the two reductions for AssemblyAI's own models on its own test sets, and the same page ranks its realtime model first on several entity measures. Treat the percentages as the vendor's claim for its pipeline. They do not transfer to another service, because a boost changes the decoder, while a post-correction map only edits text after the fact.

Word boost claims and the Sume equivalent, AssemblyAI benchmarks and Sume STT schema, read 2026-10-05
FeatureAssemblyAI (vendor claim)Sume STT 1.0
NamesEntity error -60.9% with word boostNo boost field
Technical termsEntity error -72.8% with word boostNo boost field
Where it actsDuring recognitionAfter recognition, in your code
Request fieldsVendor-specificaudio_url, language_code, duration_seconds, segmentation, delivery

A term map in 20 lines

Keep a list of the exact spellings you want: people, brands, product names. For every word in the transcript, find the closest term by string similarity above a cutoff, and swap it in. Use a high cutoff so ordinary words are left alone, and run it only on text, never on the word timings, so timing stays untouched.

import difflib, re

TERMS = ["Sume", "Cartesia", "Sonic", "Kubernetes", "Priya Raghavan"]
SINGLE = {t.lower(): t for t in TERMS if " " not in t}

def fix(text, cutoff=0.84):
    def swap(m):
        w = m.group(0)
        hit = difflib.get_close_matches(w.lower(), SINGLE, n=1, cutoff=cutoff)
        return SINGLE[hit[0]] if hit else w
    out = re.sub(r"[A-Za-z]{4,}", swap, text)
    for t in TERMS:
        if " " in t:
            out = re.sub(re.escape(t), t, out, flags=re.I)
    return out

print(fix("We call Cartesya Sonick from Kubernetes."))

Measure before you trust it

A correction map can introduce errors, by turning a real word into a similar name. Count both directions on a held-out set of ten clips: terms fixed and ordinary words broken. Raise the cutoff until broken words reach zero. The point of the exercise is the same as the vendor's: lower the error on the strings that matter most to your viewers.

At $0.01 per audio minute, ten 1-minute clips cost 10 cents, so the test is cheap.

Where the term list comes from

Start from sources you already own: the script you recorded against, the product catalogue, the guest list and the glossary in your brand guide. Add the ten words your transcripts get wrong most often, found by reading a sample of real output. Keep the list short; a map of 40 well-chosen terms beats a map of 4,000 scraped words, because every extra term adds a chance of a false swap.

Store the map in version control and note the date each term was added. If a vendor later ships a vocabulary field, the same list is what you would hand to it.

Limits of the approach

A string-similarity map fixes spelling variants of a known term. It cannot fix a name the recogniser split into two ordinary words, or a number heard wrongly. For those, keep the original audio and word timings, which Sume always returns, and review the flagged spans by hand. The map is a clean-up step, not a recogniser.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume