Sume STT language_code: set a hint or omit it for auto-detect?
language_code is optional on Sume STT. Omit it to auto-detect, pass a BCP-47 hint like en or ko when you know the language. A quick way to choose.

Pass language_code when you know the language of the recording, and omit it when you do not. On Sume STT the field is optional: the spec describes it as a BCP-47 or provider language hint, for example en or ko, and says to omit it for auto-detect. It takes one value. Sume does not publish a measured accuracy difference between hinting and auto-detect, so run a short sample of your own audio before you decide for a whole archive.
The practical rule is about what you know, not what you hope.
What does the spec say?
In the Sume OpenAPI spec, language_code is an optional hint with a single value, and omitting it asks for auto-detection. It sits beside duration_seconds, segmentation and metadata, none of which depend on it.
| Your audio | Suggested setting |
|---|---|
| One known language across the batch | Pass it, for example en or ko |
| Unknown or mixed sources | Omit and check the output |
| Two languages in one file | One hint only; split the file if the second language matters |
| Short clips with names or numbers | Pass the hint and spot check |
How do you test it cheaply?
STT is $0.01 per audio minute, so a comparison on a 2 minute clip with and without the hint costs about two cents each. Run both, read both, and keep whichever is closer to what you hear.
- Use a clip with names, numbers and fast speech, the hard parts.
- Compare transcripts by eye on the same ten sentences.
- Record which setting you chose in your own
metadata.
What does the request look like?
Pass the hint from the environment, or leave it out:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
body = {"audio_url": os.environ["AUDIO_URL"], "duration_seconds": 120}
if os.environ.get("LANG_HINT"):
body["language_code"] = os.environ["LANG_HINT"]
r = requests.post("https://api.sume.com/v1/stt-1.0/transcribe",
headers=H, json=body, timeout=60)
r.raise_for_status()
print(r.json()["data"]["job"]["id"])
What about multi-language audio?
Sume takes a single hint, so see how another vendor handles several in multiple expected languages vs Sume's single language_code. For pricing across providers, see the cheapest speech-to-text API per hour.
Sources
Related posts
More in Developers
- Sume STT takes 10 minutes per request: a 90-minute file is 9 jobs
Sume speech-to-text caps a request at 600 seconds of audio, so a 90-minute recording is nine requests at up to $0.10 each, $0.90 in total.
- Turn Sume STT word timings into an SRT file in 30 lines of Python
Sume speech-to-text returns words[] with start and end seconds. Group them into SRT cues of 42 characters or 5 seconds, and see what the caption endpoint takes.
- Count Sume job replays with idempotency_hit after a client retry
A Sume job submit returns data.idempotency_hit. Log it after a retried video submit to prove the second call replayed the first job, not a new one.
- Sume submit budgets: 120 to 1,200 writes a minute for an Omni batch
Sume rate limits submits per plan: Free 120, Pro 300, Startup 600 and Scale 1,200 a minute; reads get 40 times that. Why queue size, not the rate, paces Omni.
Written by Sume