Detect a clip's language before dubbing with Sume STT

Run Sume STT with no language hint and read language_code and language_probability from the result. A short Python check that gates a dub on a confident answer.

5 min readSume
All posts

Submit the clip to POST /v1/stt-1.0/transcribe without a language_code. Sume then detects the language, and the finished result carries language_code and language_probability next to the text. Accept the answer only when the probability is high enough for your purpose, and send the clip to a human when it is not.

This matters before dubbing because every later step depends on the source language: the translation, the target voice and the language you send to text-to-speech. A wrong guess here is the cheapest mistake to catch, since a 20-second probe costs about a third of a cent at 1 cent a minute.

What the request and result carry

The request takes a public HTTPS audio_url, an optional language_code hint and an optional duration_seconds from 1 to 600. Leave the hint out for detection. When you do send a hint, the result returns the requested code, so a result with a hint tells you nothing new about the clip.

Both result fields are optional: the schema marks them as present when available. So treat a missing language_code as no answer and not as English. The code below returns None for a missing or low-confidence answer, and gives the probability back so you can log it.

A Python probe

The script reads the key from SUME_API_KEY and fails clearly when it is not set. It uses a fresh Idempotency-Key per run, polls with a fixed 3-second wait, and returns the code only when the probability reaches your floor of 0.8. The floor is a starting point to tune against your own clips, not a Sume recommendation.

Pass the real clip length as duration_seconds, since without it Sume reserves one minute for the job.

import os, time, uuid, requests

BASE = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def detect_language(audio_url, seconds, floor=0.8):
    body = {"audio_url": audio_url, "duration_seconds": seconds}
    r = requests.post(BASE + "/v1/stt-1.0/transcribe", json=body,
                      headers={**H, "Idempotency-Key": str(uuid.uuid4())})
    r.raise_for_status()
    j = r.json()
    job = j.get("request_id") or j["id"]
    while True:
        s = requests.get(f"{BASE}/v1/jobs/{job}/status", headers=H).json()
        if s.get("status") in ("completed", "failed", "canceled"):
            break
        time.sleep(3)
    if s.get("status") != "completed":
        return None, 0.0
    res = requests.get(f"{BASE}/v1/jobs/{job}/result", headers=H).json()
    code = res.get("language_code")
    prob = res.get("language_probability") or 0.0
    return (code if code and prob >= floor else None), prob

if __name__ == "__main__":
    print(detect_language("https://example.com/clip.mp3", 20))

What to do with the answer

Use the answer to route the next step. A confident code goes to the translation step and then to Sume TTS with the target language and a voice that fits it. A low-confidence or empty answer goes to a review queue. A clip with two languages, such as an interview with an interpreter, may come back with one dominant code, so check long clips by reading the transcript.

Speech with little content gives a weak answer. A clip of music with a few spoken words, or one with long silence, may return a low probability. Trim to the stretch where someone speaks before the probe, or accept a manual language tag for those clips.

Cost and where it fits

A language probe costs a fraction of a dub. The probe is at most 10 cents for a 10-minute file, and the TTS in the target language is $0.0475 per 1,000 characters at list. Microsoft lists 23 languages for MAI-Voice-2.1 (Microsoft AI, read 2026-10-05), which is a reminder to check that your target language has a voice before you spend on the transcript.

Keep the probe's job id and probability with the clip record. When a dub sounds wrong, that record says whether the source language was detected or assumed.

Calibrate the floor

Run the probe on ten clips of known language, including two with background music. Check the codes and the probabilities, then set your floor from what you see. If every correct answer is above 0.9 and every wrong one below 0.6, a floor of 0.8 is fine, and if they overlap, send those clips to people.

Re-run on a clip only when the audio changes. Detection on the same file with the same request gives you nothing new, and a repeated submit with the same Idempotency-Key returns the same job instead of billing again.

Probe only the first 20 to 30 seconds of speech when the clip is long. Detection does not need the whole file, and a short cut keeps the probe near its minimum cost. Cut it with Timeline audio or audio detach if the source is a Sume-hosted file, or host a short copy at your own HTTPS URL.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume