Route low-confidence transcripts to review: STT language_probability

Sume STT returns language_code and language_probability. Flag results under a threshold you set and send them to a person. Python, about 10 cents per file.

4 min readSume
All posts

A batch of recorded calls comes in from five countries. Most transcribe cleanly. A few are crosstalk, music, or a language you did not expect. Reading all of them is a waste of time, and trusting all of them is a risk. Sume STT 1.0 returns two fields that let you sort: language_code and language_probability, the provider's detection confidence for the language.

What the fields mean

In the Sume job result schema, STT results return text, optional language fields and words[]. language_probability is a number when the provider supplied it and can be null. It measures confidence about the language, not accuracy of the words. A clean recording in the right language can still have errors, and a low score is a good reason to read, not proof that the text is wrong.

A review rule that is simple to defend

Pick a threshold, such as 0.8, from a sample of 20 files you have already checked by hand. Send anything below it, anything null, and any language outside your expected set to the review pile. Keep the rest. If you already know the language, pass language_code and the detection step is moot, as the language hint post explains.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

EXPECTED = {"en", "es", "fr"}
review, ok = [], []
for url in open("urls.txt").read().split():
    res = run("/stt-1.0/transcribe", {"audio_url": url, "duration_seconds": 600})
    p = res.get("language_probability")
    bad = p is None or p < 0.8 or res.get("language_code") not in EXPECTED
    (review if bad else ok).append((url, res.get("language_code"), p))
print(len(ok), "ok;", len(review), "to review")
for row in review:
    print(row)

Cost

Each STT minute is about $0.01, so a full 10-minute file is about 10 cents, and a batch of 100 files is about $10 in total. The review rule does not add any cost. It only decides who reads the text. duration_seconds is a reservation hint from 1 to 600; leave it out and one minute is reserved.

Pitfalls

  • A file with two languages gets one code. Split it at the switch, or review it.
  • Music or silence at the start can pull the score down. The score still tells you something is unusual.
  • Do not use the number as a quality score between providers. Compare only with your own hand-checked sample.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume