Transcribe a clip when you do not know the language: STT auto-detect

Omit language_code on Sume speech-to-text and the job auto-detects the language. What the result returns, the 10-minute cap, and the $0.01 a minute rate.

4 min readSume
All posts

To transcribe audio in an unknown language with Sume, call POST /v1/stt-1.0/transcribe and leave out language_code. The job auto-detects the language and the result reports a language_code, plus text and word timings (API reference). The rate is $0.01 per audio minute and one job takes at most 10 minutes of audio.

Microsoft's new MAI-Transcribe-2-Streaming covers 60 languages with automatic continuous language detection and first partial results in just over 100 ms, at $0.54 per hour of audio through the end of 2026 (Microsoft AI, read 2026-10-04). That is a live-stream product. Sume STT is a batch job for a finished file, which is the shape most short-video work needs.

What you send

Required is audio_url, a public HTTPS URL. Send duration_seconds between 1 and 600. If you omit it, Sume reserves one minute, so a 9-minute file without the hint reserves less than it needs. diarize and tag_audio_events are rejected, so expect no speaker labels.

A run you can copy

The run helper submits with an Idempotency-Key, polls the job until it is terminal, then reads the result.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, json=body, timeout=60,
                      headers={**H, "Idempotency-Key": key})
    r.raise_for_status()
    job = r.json()["request_id"]
    while True:
        s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds") or 3)
    res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
    res.raise_for_status()
    return res.json()

res = run("/v1/stt-1.0/transcribe", {
    "audio_url": os.environ["CLIP_URL"],
    "duration_seconds": 120,
}, "detect-lang-001")
print(res.get("language_code"), res.get("text", "")[:200])

What the result does and does not promise

The result carries one language_code. The docs do not describe a clip that switches language partway, so treat that as a case to test on your own audio. Microsoft's page describes continuous detection for its streaming model, and the Sume docs describe no equivalent.

If you know the language, send it. A hint costs nothing and removes one source of error. If you are sorting a folder of clips, run the detect pass on each, group by language_code, then re-run the groups you care about with the hint set.

Cost of a sorting pass

At $0.01 a minute, 100 clips of 90 seconds each is 150 minutes, or $1.50. A reserved minute per clip is the floor when you do not know the duration, so measure duration locally first and send it. The rate is the same for every language. For many files see batch transcription.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume