Detect a clip's language before dubbing with Sume STT
Run Sume STT with no language hint and read language_code and language_probability from the result. A short Python check that gates a dub on a confident answer.

Submit the clip to POST /v1/stt-1.0/transcribe without a language_code. Sume then detects the language, and the finished result carries language_code and language_probability next to the text. Accept the answer only when the probability is high enough for your purpose, and send the clip to a human when it is not.
This matters before dubbing because every later step depends on the source language: the translation, the target voice and the language you send to text-to-speech. A wrong guess here is the cheapest mistake to catch, since a 20-second probe costs about a third of a cent at 1 cent a minute.
What the request and result carry
The request takes a public HTTPS audio_url, an optional language_code hint and an optional duration_seconds from 1 to 600. Leave the hint out for detection. When you do send a hint, the result returns the requested code, so a result with a hint tells you nothing new about the clip.
Both result fields are optional: the schema marks them as present when available. So treat a missing language_code as no answer and not as English. The code below returns None for a missing or low-confidence answer, and gives the probability back so you can log it.
A Python probe
The script reads the key from SUME_API_KEY and fails clearly when it is not set. It uses a fresh Idempotency-Key per run, polls with a fixed 3-second wait, and returns the code only when the probability reaches your floor of 0.8. The floor is a starting point to tune against your own clips, not a Sume recommendation.
Pass the real clip length as duration_seconds, since without it Sume reserves one minute for the job.
import os, time, uuid, requests
BASE = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def detect_language(audio_url, seconds, floor=0.8):
body = {"audio_url": audio_url, "duration_seconds": seconds}
r = requests.post(BASE + "/v1/stt-1.0/transcribe", json=body,
headers={**H, "Idempotency-Key": str(uuid.uuid4())})
r.raise_for_status()
j = r.json()
job = j.get("request_id") or j["id"]
while True:
s = requests.get(f"{BASE}/v1/jobs/{job}/status", headers=H).json()
if s.get("status") in ("completed", "failed", "canceled"):
break
time.sleep(3)
if s.get("status") != "completed":
return None, 0.0
res = requests.get(f"{BASE}/v1/jobs/{job}/result", headers=H).json()
code = res.get("language_code")
prob = res.get("language_probability") or 0.0
return (code if code and prob >= floor else None), prob
if __name__ == "__main__":
print(detect_language("https://example.com/clip.mp3", 20))What to do with the answer
Use the answer to route the next step. A confident code goes to the translation step and then to Sume TTS with the target language and a voice that fits it. A low-confidence or empty answer goes to a review queue. A clip with two languages, such as an interview with an interpreter, may come back with one dominant code, so check long clips by reading the transcript.
Speech with little content gives a weak answer. A clip of music with a few spoken words, or one with long silence, may return a low probability. Trim to the stretch where someone speaks before the probe, or accept a manual language tag for those clips.
Cost and where it fits
A language probe costs a fraction of a dub. The probe is at most 10 cents for a 10-minute file, and the TTS in the target language is $0.0475 per 1,000 characters at list. Microsoft lists 23 languages for MAI-Voice-2.1 (Microsoft AI, read 2026-10-05), which is a reminder to check that your target language has a voice before you spend on the transcript.
Keep the probe's job id and probability with the clip record. When a dub sounds wrong, that record says whether the source language was detected or assumed.
Calibrate the floor
Run the probe on ten clips of known language, including two with background music. Check the codes and the probabilities, then set your floor from what you see. If every correct answer is above 0.9 and every wrong one below 0.6, a floor of 0.8 is fine, and if they overlap, send those clips to people.
Re-run on a clip only when the audio changes. Detection on the same file with the same request gives you nothing new, and a repeated submit with the same Idempotency-Key returns the same job instead of billing again.
Probe only the first 20 to 30 seconds of speech when the clip is long. Detection does not need the whole file, and a short cut keeps the probe near its minimum cost. Cut it with Timeline audio or audio detach if the source is a Sume-hosted file, or host a short copy at your own HTTPS URL.
Sources
Related posts
More in Developers
- Detect Sume API changes in CI: diff the live OpenAPI document
Pull https://api.sume.com/reference/json in CI, reduce it to sorted method and path lines with jq, and fail the build when the route list changes.
- Idempotency-Key from a sha256 of the body: Seedance 2.5 retries
Retry a timed-out POST /v1/videos without paying twice by deriving Idempotency-Key from the request body. Python code plus the 409 conflict rule on Sume.
- Burned-in captions and YouTube AI disclosure: captions are exempt
YouTube's disclosure page names caption creation among edits that need no altered-content label. What that covers when Sume burns captions into a Short.
- Does the previous agent turn help transcription? AssemblyAI says 4.3%
AssemblyAI reports agent context cut entity error from 16.04% to 15.35%, a 4.3% relative drop. Sume STT has no context field; here is a fair test.
Written by Sume