Thai, Vietnamese, Indonesian speech to text: Sume STT language hints
Send th, vi or id as language_code to Sume STT, or omit it to auto-detect, then check language_probability. $0.01 per audio minute; test a sample first.

To transcribe Thai, Vietnamese or Indonesian speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to th as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Thai, Vietnamese or Indonesian clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Thai, Vietnamese or Indonesian. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Use th, vi or id as the hint for Thai, Vietnamese or Indonesian. Sume does not publish a language list for its STT, so treat each as a request to test. If a short clip comes back with a different language_code, drop the hint and compare.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-th-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "th",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Thai is written without spaces between words, Vietnamese uses Latin letters with many tone marks, and Indonesian is plain Latin script. That means the same check does not fit all three. For Thai, look at how words[] breaks. For Vietnamese, make sure the tone marks survive in text. For Indonesian, a short clip may be mistaken for Malay, so read language_probability.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as th; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Thai, Vietnamese or Indonesian speech
Vietnamese and Indonesian use Latin letters, which falls inside the Latin font coverage described in the captions docs, so a burned-in test is reasonable for them. Thai is a different script, so do not plan a Thai caption render until you have tested it. Pass the verified transcript to video captions as script_text to keep the timings.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- Timeline output.fps: why a 24 fps clip judders when you force 30
Leave output.fps unset and Timeline renders at the source rate. Force 30 on a 24 fps clip and frames repeat; the job reports output_fps_resamples_sources.
- Transcribe audio with curl and jq: a Sume STT shell script
A 16-line bash script that submits audio to Sume STT, polls the job with curl, and prints every word with start and end times through jq. One cent per minute.
- unsupported_capability names sume/auto: fix a Sume ad clip request
A 400 unsupported_capability on sume/auto hides the resolved model but lists accepted values in supported. How to fix duration, resolution or audio.
- verifyWebhook in a fetch handler: four rules, 204 for unknown events
Use @sume-com/sdk verifyWebhook on the raw body, await it, treat false as 401 and answer unknown events with 204. A runnable handler for Workers, Deno and Node.
Written by Sume