Japanese speech to text API: Sume STT with language_code ja
Transcribe Japanese audio with Sume STT: send language_code ja, read word times, and test a sample first. $0.01 per audio minute, 10 minute jobs.

To transcribe Japanese speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to ja as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Japanese clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Japanese. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Set language_code to ja for Japanese. The field accepts a string of 2 to 16 characters, so a regional tag also fits the schema, but Sume does not document which tags change the output. Start with the plain ja hint and compare it against a run with the field omitted.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-ja-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "ja",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Japanese is written without spaces between words, so check how the words[] array breaks your sample before you build anything on word counts. Print the first twenty entries with their start and end values. If entries are whole phrases, time captions from segments[] or from the words as returned rather than counting characters or splitting on spaces.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as ja; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Japanese speech
For on-screen captions, the Sume captions endpoint documents Latin and Hangul font coverage only, so Japanese burned-in captions are a separate question, covered in what Sume documents for Japanese, Chinese and Arabic captions. A safe route today is to use the STT result to produce a subtitle file or a transcript and show it in your player, and burn captions only after you confirm font support.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- Job id or run id? Which Sume endpoint to poll for each product
Jobs, Format runs, Actions and Agent Completions have different ids, poll URLs and webhook events. Which to poll for each product, and which SDK helper to call.
- Let browsers start Sume jobs through your server, not with your key
Browsers must never hold a Sume API key. A server route authenticates the user, checks the input, derives an Idempotency-Key and returns only the status URL.
- How do I add a listen-to-this-page audio version with TTS?
Turn each article into an audio file with one async TTS job per page: a 9,000-character article costs 43 cents on Sume. What it does not replace.
- Mandarin Chinese speech to text API: Sume STT language_code zh
Transcribe Mandarin audio with Sume STT using language_code zh, then check the result and timings. $0.01 per audio minute and no accuracy claim without a test.
Written by Sume