Moroccan Arabic speech to text: open-source Voxtral model vs Sume STT

A Moroccan open-source stack pairs language ID with Voxtral ASR for dialect code-switching. What Sume STT offers for Arabic audio, and what it lacks.

4 min readSume
All posts

For Moroccan Arabic with French and English mixed in, the new open-source effort is the better-aimed tool. TechAfrica News reported on October 5 that Morocco released an open-source language-identification classifier and a Voxtral-based speech recognizer, built with Mistral, for Moroccan dialect and Arabic, French and English code-switching. Sume STT can transcribe Arabic audio, but it has no dialect setting and no published Darija accuracy.

That is the honest comparison. A specialised open model may handle a dialect better; a hosted job is quicker to wire up. Decide with a test on your own clips, not with a headline.

What each option states

The Morocco column comes from the TechAfrica News report. The Sume column comes from the API schema.

Dialect speech options (read 2026-10-08)
ItemMorocco open-source stackSume STT
TargetMoroccan dialect, code-switchingGeneral multilingual audio
PartsLanguage-ID classifier plus Voxtral ASROne hosted endpoint
HostingYou run it (open source)Sume runs it
Language inputClassifier decidesOptional language_code hint, or auto-detect
Dialect controlBuilt for itNone
CostYour computeAbout $0.01 per minute

Test Sume STT on dialect clips in four steps

  • Collect 20 clips of real Darija speech, each under 10 minutes.
  • Submit each with language_code set to ar, then again with the field omitted.
  • Compare both transcripts with a human-checked reference.
  • Count the errors on names and numbers, not just the overall rate.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/stt-1.0/transcribe",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "stt-demo-001",
    },
    json={
        "audio_url": "https://media.sume.com/example/clip.mp3",
        "duration_seconds": 540,
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

Why dialect is a separate problem

General speech models are trained mostly on standard forms of a language. A dialect that mixes Arabic with French and English inside one sentence can trip a model that expects a single language per clip. A language-identification step that runs before recognition, as the Moroccan stack describes, is aimed at exactly that case.

With a hosted endpoint you can only choose the hint. If the hint is ar and the speaker switches to French mid-sentence, you cannot tell the engine to expect it. Keep clips short and label results by speaker turn on your side if you need to measure the switch points.

What Sume does not do

Sume STT does not select a regional dialect and does not promise code-switching quality between Arabic and French. The language_code field is a hint, and omitting it lets the engine detect the language. Our Arabic STT guide covers the hint in more detail. The TechAfrica article is a news report; I did not verify accuracy figures for the Moroccan models.

Next step

Run the measurement script from measure word error rate on your own clips on both systems. If the open model wins on dialect clips, use it. If the hosted job is close, the lower setup cost may decide. Keep the reference transcripts you write for this test; they let you re-score any model that launches next month in minutes, instead of starting the evaluation again from nothing. Record the date of each run and the exact settings used, since both change over time.

Sources

Related posts

More in Models

All Models posts

Written by Sume