Arabic speech to text API: Sume STT with language_code ar
Transcribe Arabic audio with Sume STT: send language_code ar, check the reported language, and review the text. $0.01 per audio minute.

To transcribe Arabic speech with Sume, submit the audio to POST /v1/stt-1.0/transcribe with language_code set to ar as a hint, wait for the job to complete, and read text and words[] from the result. Sume prices STT at $0.01 per audio minute, so a 90 second Arabic clip reserves and bills about one and a half cents of audio time when you pass duration_seconds.
One honest caveat first. The Sume docs do not publish a list of supported languages for STT, so this page does not promise accuracy for Arabic. It shows how to send the hint, how to check what came back, and how to judge a sample before you commit to a batch.
The request
Use ar as the hint for Arabic. Spoken Arabic varies a lot by region, and the Sume docs give no dialect guidance, so treat the hint as a starting point and check the result against a sample from your own speakers. Try the clip with and without the hint.
The job is asynchronous. Submit returns 202 with a request_id, which is the job id. Poll GET /v1/jobs/{id}/status until it is completed, failed or canceled, then read GET /v1/jobs/{id}/result. Always send an Idempotency-Key so a retried submit does not create a second paid job.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-ar-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"language_code": "ar",
"duration_seconds": 90,
"segmentation": "sentence"
}'
# then poll GET /v1/jobs/$REQUEST_ID/status and read /resultCheck the result before you trust it
Arabic text runs right to left, so check how the transcript renders where you store it. The text is a plain string and the words[] entries are in speech order with start and end seconds. Keep those timings as numbers, and only reverse or shape text in the display layer, not in the data you store.
Compare the language_code in the result with the one you sent, and read language_probability. If they disagree, the clip may be mixed-language, very short, or mostly music. Run two or three of your own recordings, including a noisy one, and read the text yourself or have a speaker of the language read it.
| Field or fact | Where | Why it matters |
|---|---|---|
| language_code (request) | POST body, 2 to 16 characters | A hint such as ar; omit it to let the service detect |
| language_code (result) | GET /v1/jobs/{id}/result | What the job reports it heard |
| language_probability | Same result | Low values mean review the text by hand |
| words[] | Same result | word, start, end in seconds for timing |
| Model page language count | microsoft.ai MAI-Transcribe-2 | 60 listed for that vendor model, not for Sume |
Captions from Arabic speech
Burned-in Arabic captions are not something the Sume captions docs cover: font coverage is documented for Latin and Hangul only, as explained in the stored note on Japanese, Chinese and Arabic captions. Use the transcript and word times for your own player's subtitle track, and verify with a short test before you plan a caption render.
Cost and limits
One STT job takes at most 600 seconds of audio. For a longer recording, split it into slices, send one job per slice, and add each slice's start time to its word times when you merge them. Splitting a video's audio first with audio detach costs $0.01 per job and gives you a 16 kHz mono WAV.
If you leave duration_seconds out, Sume reserves one minute, which is wrong for a long file. Pass the real length rounded up. Microsoft's current model page lists 60 languages for its transcription model (read 2026-10-07), which is a vendor claim about that model and not a statement about Sume's STT.
The result also has segments[] when you ask for segmentation, and each segment carries start and end seconds. For subtitles you do not need to cut the text yourself: pass the transcript or your edited script to the captions endpoint and let it keep the speech timings.
Sources
Related posts
More in Developers
- Avatar video webhook mode: what arrives and what to poll anyway
Use mode webhook for a Sume avatar video and Sume posts one terminal event: completed, failed or canceled. Payload, signature headers and the polling backup.
- Balance check before an ad variant burst: Sume API 402 guard
Read GET /v1/balance before sending a burst of video variants. Sume reserves 1.25 times list price on submit and returns 402 insufficient_credits below it.
- How do I make a bilingual English and Spanish audio announcement?
Make one bilingual announcement file: two TTS jobs, one per language, joined by a $0.01 Timeline audio concat with no re-synthesis. About 10 cents in total.
- callback_url or webhook_url: which field each Sume video route takes
POST /v1/videos takes callback_url; motion control, lip-sync and image routes take mode plus webhook_url. The field names and what they share.
Written by Sume