Mandarin Chinese speech to text API: Sume STT language_code zh
Transcribe Mandarin audio with Sume STT using language_code zh, then check the result and timings. $0.01 per audio minute and no accuracy claim without a test.

To transcribe Mandarin Chinese speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to zh as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Mandarin Chinese clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Mandarin Chinese. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Use zh as the hint for Chinese. The schema allows up to 16 characters, so zh-CN also fits, but Sume does not document which regional tags change the output or whether Simplified or Traditional characters come back. Run a sample and look at the script of the returned text yourself.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-zh-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "zh",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Chinese, like Japanese, has no spaces between words, so words[] may not line up with what a reader sees as words. Print a few entries with their times. If you plan to chunk text for captions, chunk by segments[] or by punctuation in text, and check the longest line against your display width.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as zh; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Mandarin Chinese speech
As with other non-Latin scripts, Sume's captions docs describe font coverage for Latin and Hangul only, so use the transcript and word times for a subtitle file in your own player until you have verified a burned-in render. See what Sume documents for Japanese, Chinese and Arabic captions for what is and is not covered.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- Migrate a real-time avatar prototype to Sume async jobs: what changes
Moving from a live avatar session to Sume means replacing a stream with submit, poll and fetch. The code changes, the UX changes, and a Node example that runs.
- Model an AI generation job as a state machine in your database
A schema and update rule for tracking Sume jobs: five statuses, sticky terminal states, a separate webhook delivery column, and the idempotency key on the row.
- Pin the model id in an ad test: sume/auto follows the catalog
sume/auto is a pure function of the request plus the catalog version, so two ad arms made weeks apart can land on different models. Pin an explicit id in tests.
- Poll hundreds of AI jobs without a thundering herd: jitter and budgets
Poll many Sume jobs without synchronized bursts: jitter, next_poll_after_seconds, per-plan read budgets, and the math on how much polling a plan can absorb.
Written by Sume