Moroccan Arabic speech to text: open-source Voxtral model vs Sume STT
A Moroccan open-source stack pairs language ID with Voxtral ASR for dialect code-switching. What Sume STT offers for Arabic audio, and what it lacks.

For Moroccan Arabic with French and English mixed in, the new open-source effort is the better-aimed tool. TechAfrica News reported on October 5 that Morocco released an open-source language-identification classifier and a Voxtral-based speech recognizer, built with Mistral, for Moroccan dialect and Arabic, French and English code-switching. Sume STT can transcribe Arabic audio, but it has no dialect setting and no published Darija accuracy.
That is the honest comparison. A specialised open model may handle a dialect better; a hosted job is quicker to wire up. Decide with a test on your own clips, not with a headline.
What each option states
The Morocco column comes from the TechAfrica News report. The Sume column comes from the API schema.
| Item | Morocco open-source stack | Sume STT |
|---|---|---|
| Target | Moroccan dialect, code-switching | General multilingual audio |
| Parts | Language-ID classifier plus Voxtral ASR | One hosted endpoint |
| Hosting | You run it (open source) | Sume runs it |
| Language input | Classifier decides | Optional language_code hint, or auto-detect |
| Dialect control | Built for it | None |
| Cost | Your compute | About $0.01 per minute |
Test Sume STT on dialect clips in four steps
- Collect 20 clips of real Darija speech, each under 10 minutes.
- Submit each with language_code set to ar, then again with the field omitted.
- Compare both transcripts with a human-checked reference.
- Count the errors on names and numbers, not just the overall rate.
import os, requests
r = requests.post(
"https://api.sume.com/v1/stt-1.0/transcribe",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "stt-demo-001",
},
json={
"audio_url": "https://media.sume.com/example/clip.mp3",
"duration_seconds": 540,
},
timeout=30,
)
r.raise_for_status()
print(r.json())Why dialect is a separate problem
General speech models are trained mostly on standard forms of a language. A dialect that mixes Arabic with French and English inside one sentence can trip a model that expects a single language per clip. A language-identification step that runs before recognition, as the Moroccan stack describes, is aimed at exactly that case.
With a hosted endpoint you can only choose the hint. If the hint is ar and the speaker switches to French mid-sentence, you cannot tell the engine to expect it. Keep clips short and label results by speaker turn on your side if you need to measure the switch points.
What Sume does not do
Sume STT does not select a regional dialect and does not promise code-switching quality between Arabic and French. The language_code field is a hint, and omitting it lets the engine detect the language. Our Arabic STT guide covers the hint in more detail. The TechAfrica article is a news report; I did not verify accuracy figures for the Moroccan models.
Next step
Run the measurement script from measure word error rate on your own clips on both systems. If the open model wins on dialect clips, use it. If the hosted job is close, the lower setup cost may decide. Keep the reference transcripts you write for this test; they let you re-score any model that launches next month in minutes, instead of starting the evaluation again from nothing. Record the date of each run and the exact settings used, since both change over time.
Sources
Related posts
More in Models
- Nano Banana 2.1 or Pro: 10 vs 19 cents, and four ratios Pro lacks
On Sume, Nano Banana 2.1 bills 10 cents and Pro bills 19. Both take edits and four resolution tiers, but Pro does not list 4:1, 1:4, 8:1 or 1:8.
- Nano Banana 2.1 vs Pro from 512 to 4K: price per image on Sume
Nano Banana 2.1 runs $0.075 at 512 to $0.20 at 4K; Nano Banana Pro is $0.1875 up to 2K and $0.375 at 4K. Full tier table and a per-1,000 view.
- nano-banana-2 still works on Sume: it runs as Nano Banana 2.1 at 10c
Nano Banana 2 is retired on Sume. The ids nano-banana-2 and google/nano-banana-2 still work and run as Nano Banana 2.1, and the job stores the 2.1 id.
- Nano Banana 2 shutdown on Oct 29? Google's changelog gives no date
Some pages quote an Oct 29 shutdown for Nano Banana 2. Google's Oct 6 changelog lists no date. What that means for a Sume image call.
Written by Sume