Hindi speech to text API: Sume STT with language_code hi
Transcribe Hindi audio with Sume STT: send language_code hi, check the reported language, and review code-mixed speech. $0.01 per audio minute.

To transcribe Hindi speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to hi as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Hindi clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Hindi. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Use hi as the hint for Hindi. The field takes a string of 2 to 16 characters. Sume does not document a language list, so run a sample from your own speakers, with and without the hint, and keep the version whose text you would publish.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-hi-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "hi",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Hindi speech often mixes in English words. If your recordings do, send a sample with that mix and read the text closely: see which script each English word comes back in, and read language_probability. Decide on a rule for your catalog, such as keeping the returned text as is, and apply it to every file.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as hi; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Hindi speech
Hindi is written in Devanagari, which is not in the font coverage the captions docs describe, since they document Latin and Hangul only. Use the transcript and word times in your own subtitle track, and test a short burned-in render before you plan one. If you give the captions endpoint a script, it keeps the speech timings.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- A 3-minute Timeline render: poll job status, don't sleep a fixed time
Timeline renders are async by default. Submit with an Idempotency-Key, poll /v1/jobs/:id/status until terminal, then read /result. Cost: 3 minutes is $0.30.
- Image-to-video not starting on my photo: frame_images vs references
Your photo is a reference, not a first frame, when it goes in input_references. Use frame_images with first_frame on Sume /v1/videos to pin the opening shot.
- Japanese speech to text API: Sume STT with language_code ja
Transcribe Japanese audio with Sume STT: send language_code ja, read word times, and test a sample first. $0.01 per audio minute, 10 minute jobs.
- Job id or run id? Which Sume endpoint to poll for each product
Jobs, Format runs, Actions and Agent Completions have different ids, poll URLs and webhook events. Which to poll for each product, and which SDK helper to call.
Written by Sume