Test a language on Sume STT with a 30-second sample first

Vendors split languages into trained and verified. On Sume STT, language_code is a hint, so cut a 30-second sample and compare auto-detect with your hint.

4 min readSume
All posts

On Sume STT, language_code is an optional BCP-47 hint such as en or ko, and omitting it means auto-detect. The API reference I read lists no table of supported languages, so the only proof for your language is a test: cut a 30-second sample, transcribe it once on auto-detect and once with your hint, then compare the text.

I read the Sume API reference, the audio detach docs and the video inspect docs on 2026-10-03, plus Meta's Muse Voice Transcribe page. Meta's page gives price, speaker and keyword-biasing claims but no language list, so I do not quote any language count for Muse here.

What does language_code actually do?

The request schema describes it as a BCP-47 or provider language hint, with a length of 2 to 16 characters. It is a hint, not a guarantee that a language is supported. Completed results carry text, words[] with start and end times in seconds, and language fields when available, so you can see what was detected.

How do you cut a 30-second sample?

If your audio is a hosted video, audio detach takes a range of start and end seconds and returns a wav at $0.01 per job. Set channels to mono and sample_rate to 16000 for the speech-to-text shape the docs describe. If you already have an audio file at a public HTTPS URL, trim it yourself and skip this step.

How do you compare auto-detect with your hint?

This script sends the same sample twice, with a 30-second duration hint so the usage reservation matches. It prints the detected language, the word count and the first 120 characters. Set SAMPLE_URL to your clip and CODE to your language, for example ko.

import json, os, urllib.request

def api(method, url, body=None):
    req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)["data"]

def run(sample_url, language_code=None):
    body = {"audio_url": sample_url, "duration_seconds": 30, "mode": "sync", "wait_timeout_seconds": 30}
    if language_code:
        body["language_code"] = language_code
    job = api("POST", "https://api.sume.com/v1/stt-1.0/transcribe", body)
    result = api("GET", job["result_url"])["result"]
    return result["text"], result.get("language_code"), len(result.get("words", []))

if __name__ == "__main__":
    url, code = os.environ["SAMPLE_URL"], os.environ["CODE"]
    for label, hint in (("auto", None), (code, code)):
        text, detected, count = run(url, hint)
        print(label, detected, count, "words:", text[:120])

What counts as a pass?

Read the text with someone who speaks the language. Then score it against a human reference using the word error rate function in how to measure word error rate on your own clips. Run at least one sample per accent and recording condition you ship, and keep the reference text next to the sample so the next model can be scored the same way.

What to check on each 30-second sample, read 2026-10-03
CheckAuto-detect runHint run
language_code in the resultIs it your language?Does it match your hint?
words[] countPlausible for 30 s?Close to the auto run?
Text vs human referenceWord error rateWord error rate

Sources

Related posts

More in Developers

All Developers posts

Written by Sume