Speech to text Korean API: send language_code or auto-detect?

Sume STT 1.0 takes an optional language_code such as ko; omit it and the language is auto-detected. When to pin Korean, when to omit it, and how TTS differs.

4 min readSume
All posts

With Sume STT 1.0, you can transcribe Korean either way: send language_code: "ko" as a hint, or omit the field and the language is auto-detected. The schema calls the field an optional BCP-47 or provider language hint. Pin ko when you know the recording is Korean, and omit it when you do not know what you will receive.

This is from the STT 1.0 request schema in the OpenAPI document behind the API reference, read 2026-09-29. This post compares the request shapes and leaves quality to a test on your own audio. The wider walkthrough is in Korean speech to text API.

What does the request look like with and without a hint?

Both requests take a public HTTPS audio_url (the schema prefers a Sume media or attachment URL). The only difference is one field.

# Pinned to Korean
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav", "language_code": "ko", "duration_seconds": 90 }'

# Auto-detect: leave language_code out
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "audio_url": "https://media.sume.com/artifacts/artf_demo/call.wav" }'

When should I pin ko?

Pin it when the source is fixed: a Korean support line, a Korean podcast feed, a Korean course library. You are telling the provider what to expect instead of asking it to guess, and the request states your intent. That is a judgment on our part, not a measured result, so compare both on a handful of your own files.

Choosing between a language hint and auto-detect on STT 1.0, read 2026-09-29.
Your audioSuggested `language_code`Why
Always KoreankoThe source is fixed, so say so
Language unknown per fileOmitThe schema says omit for auto-detect
Mixed languages in one fileTest bothThe schema says nothing about code-switching

What else changes the request?

duration_seconds is optional, from 1 to 600. It only sets the usage reservation, and if you omit it Sume reserves for 1 minute. Word timings are always returned with no flag to enable them; add segmentation with mode: "sentence" to also get sentence segments.

Is text to speech the same?

No, and this trips people up. On the TTS side, the language field defaults to English at the provider when you omit it, and the schema says to set it for every non-English transcript. Sume infers ko or ja from a Hangul- or kana-only transcript as a fallback. So on STT omitting the hint means auto-detect, while on TTS omitting it means English, unless the text is purely Hangul.

What if the recording is a video?

Extract the audio first. The Audio detach API pulls a video's audio track into a Sume-hosted file, which then works as the audio_url. Extraction is a separate job, so run it once and reuse the file for every transcription you try. A related post covers audio detach for speech to text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume