Voxtral Mini 4B Realtime Arabic: open weights vs a Sume STT job

Mistral's Arabic streaming model arrived Oct 8 as Apache 2.0 weights, not a service. For Arabic transcripts without hosting, Sume STT takes language_code ar.

5 min readSume
All posts

Voxtral Mini 4B Realtime Arabic is a downloadable model, not a hosted API: its Hugging Face card lists Apache 2.0 weights, about 4.4 billion parameters, 16 kHz audio input and a configurable delay, with an average character error rate of 8.82% at 480 ms across seven Arabic benchmarks. If you want an Arabic transcript without running a GPU, Sume STT accepts an optional language_code such as ar and returns text plus word timings from a hosted job.

The model facts come from the Hugging Face model card (read 2026-10-10). Sume's request fields come from the Sume OpenAPI contract. Sume does not list Voxtral; its STT model id is sume/stt-1.0, and provider model ids stay internal.

What the model card says

The card describes a streaming transcription model for Modern Standard Arabic and dialects, fine-tuned from Voxtral Mini 4B Realtime 2602, co-developed with Morocco's digital transition ministry. It lists a causal encoder-decoder, serving through vLLM with WebSocket streaming to a /v1/realtime endpoint or through Transformers 5.2.0 or newer, and states that other languages need the original Voxtral Mini 4B Realtime model. The license is Apache 2.0 with a clause not to violate third-party rights.

Self-hosted Voxtral Arabic vs a hosted Sume STT job (card read 2026-10-10; Sume per OpenAPI)
QuestionVoxtral Mini 4B Realtime ArabicSume STT 1.0
How you run itDownload weights; serve on your own acceleratorSubmit a job to POST /v1/stt-1.0/transcribe
Input16 kHz audio streamPublic HTTPS audio_url
Language controlArabic onlyOptional language_code; omit for auto-detect
Latency shapeStreaming, configurable delay (480 ms tested)Async job; sync waits at most 30 seconds
Cost shapeYour hardware$0.01 per audio minute

What a Sume Arabic job looks like

The request needs one field, audio_url, a public HTTPS URL, preferably a Sume media URL. Add language_code: "ar" if you know the language; the OpenAPI text says to omit it for auto-detect. You can request segmentation.mode: "sentence" to get gapless sentence segments, and word timings are always returned. Provider knobs such as diarization are fixed server-side, so there is no speaker-label switch.

Send duration_seconds (1 to 600) if you know it; the docs say omitting it reserves one minute, and a single job is capped at 10 minutes. A 12-minute interview therefore needs two jobs.

Dialects, honestly

Voxtral's card names sixteen varieties and tests them on Arabic benchmarks. Sume's documentation does not list Arabic dialects, only that language_code is a BCP-47 style hint. So for a Moroccan or Gulf recording, run a short test clip on Sume and read the output before you promise a customer dialect accuracy. The existing Moroccan Arabic post covers the earlier Voxtral release.

If your recording is a video, detach the audio first: 16 kHz mono WAV is the STT shape the detach page names, and a detach job costs $0.01.

When to host it yourself

Pick the open weights if you need live captions with a sub-second delay, data that never leaves your servers, or fine-tuning. Pick the Sume job if you transcribe files, want word timings without running a server, and can accept a polled result. The two do not compete on streaming: Sume's STT route is not a streaming endpoint, and the docs describe sync as a bounded wait rather than a live feed.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume