language_code or auto-detect on Sume STT? A two-arm test on your clips

AssemblyAI reports 8.4% mean WER over 18 FLEURS languages. For your audio, run each clip twice on Sume STT, with and without language_code, and compare.

5 min readSume
All posts

Run both and keep the one that scores better on your own clips. Sume STT takes an optional language_code hint and auto-detects the language when you omit it, and neither choice is right for every recording. AssemblyAI's benchmarks page reports a mean word error rate of 8.4% across 18 FLEURS languages, and a 7.20% rate on code-switched audio across five language pairs. Those figures say language mix matters; they do not say which setting is best for your files.

Where each setting wins

A hint helps when you know the language and the audio is hard: noisy, short or heavily accented. It can hurt when a clip switches languages, because the hint pushes every word toward one language. Auto-detect is safer for mixed content and for archives where the language is unknown, but a very short clip may not give it enough to decide on.

Language-handling choices for Sume STT, Sume STT schema and AssemblyAI benchmarks page, read 2026-10-05
SituationTry firstWhy
Known language, noisy clipSet language_codeRemoves a detection step
Clip switches languagesOmit language_codeA single hint can bias every word
Unknown archiveOmit language_codeNo metadata to hint from
Very short clipSet language_code if knownLittle audio to detect from

The two-arm test

Take 10 clips and a plain-text reference for each. Transcribe each clip twice and compute the word error rate against the reference. Two arms on 10 clips of one minute is 20 jobs and 20 cents. Use en or es or whatever your language code is; the field accepts values of 2 to 16 characters.

import os, re, requests, time
H = {"x-api-key": os.environ["SUME_API_KEY"]}

def words(s): return re.findall(r"\w+", s.lower())

def wer(ref, hyp):
    r, h = words(ref), words(hyp)
    d = list(range(len(h) + 1))
    for i, rw in enumerate(r, 1):
        prev, d[0] = d[0], i
        for j, hw in enumerate(h, 1):
            prev, d[j] = d[j], min(d[j] + 1, d[j - 1] + 1, prev + (rw != hw))
    return d[-1] / max(1, len(r))

def flat(x):
    if isinstance(x, str): return x
    v = x.values() if isinstance(x, dict) else x if isinstance(x, list) else []
    return " ".join(flat(i) for i in v)

print(wer("order four seven", "order for seven"))  # 0.33

A caution about flattening

Sume always returns word timings, so a result may contain the same text in several fields. The flat helper above joins every string, which would double-count words if both a full text and a word list are present. Before you compute WER on real output, print one result, find the field that holds the full transcript, and score only that. Then choose the arm with the lower average WER, and re-run the test when your audio mix changes.

What to record from each run

Keep a row per clip and arm: clip id, whether language_code was set, the code, the word error rate, and the date. After ten clips, compute the average for each arm and the number of clips where each arm won. A tie on average with a win on the hard clips tells you to set the hint only for hard clips. Re-run the test when your audio mix changes, such as when you add a market.

Remember what the benchmark figures are. An 8.4% mean across 18 languages is an average of very different languages, and a 7.20% rate on code-switching covers five language pairs. A single language on your own recordings may sit well above or below either number.

Cost of the experiment

Ten one-minute clips in two arms is 20 jobs, 20 minutes of audio, and 20 x $0.01 = $0.20. Set duration_seconds to 60 so the reservation matches the clip length.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume