Speech to text Korean API: send language_code or auto-detect?
Sume STT 1.0 takes an optional language_code such as ko; omit it and the language is auto-detected. When to pin Korean, when to omit it, and how TTS differs.

With Sume STT 1.0, you can transcribe Korean either way: send language_code: "ko" as a hint, or omit the field and the language is auto-detected. The schema calls the field an optional BCP-47 or provider language hint. Pin ko when you know the recording is Korean, and omit it when you do not know what you will receive.
This is from the STT 1.0 request schema in the OpenAPI document behind the API reference, read 2026-09-29. This post compares the request shapes and leaves quality to a test on your own audio. The wider walkthrough is in Korean speech to text API.
What does the request look like with and without a hint?
Both requests take a public HTTPS audio_url (the schema prefers a Sume media or attachment URL). The only difference is one field.
# Pinned to Korean
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav", "language_code": "ko", "duration_seconds": 90 }'
# Auto-detect: leave language_code out
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav" }'When should I pin ko?
Pin it when the source is fixed: a Korean support line, a Korean podcast feed, a Korean course library. You are telling the provider what to expect instead of asking it to guess, and the request states your intent. That is a judgment on our part, not a measured result, so compare both on a handful of your own files.
| Your audio | Suggested `language_code` | Why |
|---|---|---|
| Always Korean | ko | The source is fixed, so say so |
| Language unknown per file | Omit | The schema says omit for auto-detect |
| Mixed languages in one file | Test both | The schema says nothing about code-switching |
What else changes the request?
duration_seconds is optional, from 1 to 600. It only sets the usage reservation, and if you omit it Sume reserves for 1 minute. Word timings are always returned with no flag to enable them; add segmentation with mode: "sentence" to also get sentence segments.
Is text to speech the same?
No, and this trips people up. On the TTS side, the language field defaults to English at the provider when you omit it, and the schema says to set it for every non-English transcript. Sume infers ko or ja from a Hangul- or kana-only transcript as a fallback. So on STT omitting the hint means auto-detect, while on TTS omitting it means English, unless the text is purely Hangul.
What if the recording is a video?
Extract the audio first. The Audio detach API pulls a video's audio track into a Sume-hosted file, which then works as the audio_url. Extraction is a separate job, so run it once and reuse the file for every transcription you try. A related post covers audio detach for speech to text.
Sources
Related posts
More in Developers
- Java subtitle generator API: burn captions with HttpClient
Generate subtitles from Java with the JDK HttpClient: POST the video URL to a captions API, poll the job, then read the captioned video_url.
- 503 provider_capacity_exceeded: same idempotency key or new?
A Sume 503 provider_capacity_exceeded means the provider queue is full. Docs say retry later with the same key; code replays the refusal, so use a new key.
- Which field do I branch on in a Sume API error: code or next_action?
Branch on the HTTP status, then on error.code. Use retryable, retry_after_seconds and next_action to decide on a resend. Never match on message.
- Which model did my AI video use? sume/auto does not say
Sume never discloses which family ran a sume/auto video: the response echoes sume/auto. What the docs say, why not to infer it, and how to pin a model instead.
Written by Sume