Voxtral Mini 4B Realtime Arabic: open weights vs a Sume STT job
Mistral's Arabic streaming model arrived Oct 8 as Apache 2.0 weights, not a service. For Arabic transcripts without hosting, Sume STT takes language_code ar.

Voxtral Mini 4B Realtime Arabic is a downloadable model, not a hosted API: its Hugging Face card lists Apache 2.0 weights, about 4.4 billion parameters, 16 kHz audio input and a configurable delay, with an average character error rate of 8.82% at 480 ms across seven Arabic benchmarks. If you want an Arabic transcript without running a GPU, Sume STT accepts an optional language_code such as ar and returns text plus word timings from a hosted job.
The model facts come from the Hugging Face model card (read 2026-10-10). Sume's request fields come from the Sume OpenAPI contract. Sume does not list Voxtral; its STT model id is sume/stt-1.0, and provider model ids stay internal.
What the model card says
The card describes a streaming transcription model for Modern Standard Arabic and dialects, fine-tuned from Voxtral Mini 4B Realtime 2602, co-developed with Morocco's digital transition ministry. It lists a causal encoder-decoder, serving through vLLM with WebSocket streaming to a /v1/realtime endpoint or through Transformers 5.2.0 or newer, and states that other languages need the original Voxtral Mini 4B Realtime model. The license is Apache 2.0 with a clause not to violate third-party rights.
| Question | Voxtral Mini 4B Realtime Arabic | Sume STT 1.0 |
|---|---|---|
| How you run it | Download weights; serve on your own accelerator | Submit a job to POST /v1/stt-1.0/transcribe |
| Input | 16 kHz audio stream | Public HTTPS audio_url |
| Language control | Arabic only | Optional language_code; omit for auto-detect |
| Latency shape | Streaming, configurable delay (480 ms tested) | Async job; sync waits at most 30 seconds |
| Cost shape | Your hardware | $0.01 per audio minute |
What a Sume Arabic job looks like
The request needs one field, audio_url, a public HTTPS URL, preferably a Sume media URL. Add language_code: "ar" if you know the language; the OpenAPI text says to omit it for auto-detect. You can request segmentation.mode: "sentence" to get gapless sentence segments, and word timings are always returned. Provider knobs such as diarization are fixed server-side, so there is no speaker-label switch.
Send duration_seconds (1 to 600) if you know it; the docs say omitting it reserves one minute, and a single job is capped at 10 minutes. A 12-minute interview therefore needs two jobs.
Dialects, honestly
Voxtral's card names sixteen varieties and tests them on Arabic benchmarks. Sume's documentation does not list Arabic dialects, only that language_code is a BCP-47 style hint. So for a Moroccan or Gulf recording, run a short test clip on Sume and read the output before you promise a customer dialect accuracy. The existing Moroccan Arabic post covers the earlier Voxtral release.
If your recording is a video, detach the audio first: 16 kHz mono WAV is the STT shape the detach page names, and a detach job costs $0.01.
When to host it yourself
Pick the open weights if you need live captions with a sub-second delay, data that never leaves your servers, or fine-tuning. Pick the Sume job if you transcribe files, want word timings without running a server, and can accept a polled result. The two do not compete on streaming: Sume's STT route is not a streaming endpoint, and the docs describe sync as a bounded wait rather than a live feed.
Sources
Related posts
More in Comparisons
- Wan 3.0 edit and extension at Alibaba vs Sume's Wan row
Alibaba lists Wan 3.0 editing and extension; Sume's wan-3.0 row does text, image, end frame and references only. Edit with Omni Flash 1.1 or H3 Max Recast.
- Wan 3.0 vs Seedance 2.5: the keep rate where Wan is cheaper
Wan 3.0 at 720p costs $0.63 per 5 s and Seedance 2.5 costs $2.89 on Sume. Wan wins per kept clip unless its keep rate falls below about 21.8% of Seedance's.
- Will Meta label a Sume video ad as AI? What Meta says it detects
Meta says it will detect third-party AI ads via industry-standard signals and add AI info to About this ad. The Sume docs promise nothing either way.
- Grok Imagine mid-video frame pin vs Sume's first and last frames
xAI lets Grok Imagine 1.5 pin first, last and mid frames. Sume's Grok row takes one start still; this maps pinning to Sume rows with first and end frames.
Written by Sume