Transcribe a clip when you do not know the language: STT auto-detect
Omit language_code on Sume speech-to-text and the job auto-detects the language. What the result returns, the 10-minute cap, and the $0.01 a minute rate.

To transcribe audio in an unknown language with Sume, call POST /v1/stt-1.0/transcribe and leave out language_code. The job auto-detects the language and the result reports a language_code, plus text and word timings (API reference). The rate is $0.01 per audio minute and one job takes at most 10 minutes of audio.
Microsoft's new MAI-Transcribe-2-Streaming covers 60 languages with automatic continuous language detection and first partial results in just over 100 ms, at $0.54 per hour of audio through the end of 2026 (Microsoft AI, read 2026-10-04). That is a live-stream product. Sume STT is a batch job for a finished file, which is the shape most short-video work needs.
What you send
Required is audio_url, a public HTTPS URL. Send duration_seconds between 1 and 600. If you omit it, Sume reserves one minute, so a 9-minute file without the hint reserves less than it needs. diarize and tag_audio_events are rejected, so expect no speaker labels.
A run you can copy
The run helper submits with an Idempotency-Key, polls the job until it is terminal, then reads the result.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job = r.json()["request_id"]
while True:
s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
if s.get("terminal"):
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
res.raise_for_status()
return res.json()
res = run("/v1/stt-1.0/transcribe", {
"audio_url": os.environ["CLIP_URL"],
"duration_seconds": 120,
}, "detect-lang-001")
print(res.get("language_code"), res.get("text", "")[:200])What the result does and does not promise
The result carries one language_code. The docs do not describe a clip that switches language partway, so treat that as a case to test on your own audio. Microsoft's page describes continuous detection for its streaming model, and the Sume docs describe no equivalent.
If you know the language, send it. A hint costs nothing and removes one source of error. If you are sorting a folder of clips, run the detect pass on each, group by language_code, then re-run the groups you care about with the hint set.
Cost of a sorting pass
At $0.01 a minute, 100 clips of 90 seconds each is 150 minutes, or $1.50. A reserved minute per clip is the floor when you do not know the duration, so measure duration locally first and send it. The rate is the same for every language. For many files see batch transcription.
Sources
Related posts
More in Developers
- Transcribe audio from a private bucket: Sume STT needs a public URL
Sume STT takes a public HTTPS audio_url. If your files sit in a private bucket, here is how to hand them over without opening the whole bucket.
- Transcribe two minutes of a long video: audio detach range, then STT
Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.
- 150-language subtitles: which scripts Sume captions document
A translation model can output 150 languages, but Sume documents Latin and Hangul caption styles. Test other scripts on a short clip before a batch.
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
Written by Sume