Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists

Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.

4 min readSume
All posts

Mistral's model overview lists Voxtral Mini Transcribe 2 and Voxtral Realtime, both versioned v26.02, alongside Voxtral TTS v26.03 with zero-shot voice cloning. The page is the place to start for speech-to-text with Mistral, but it gives little else about limits, so confirm the details on each model's own page before you build.

What the overview lists

The overview groups the audio models by task. The transcription entries cover a batch model and a realtime model, and the speech entry covers synthesis with zero-shot voice cloning.

Voxtral models on Mistral's overview page (read 2026-10-03)
ModelVersionListed capability
Voxtral Mini Transcribe 2v26.02Speech-to-text
Voxtral Realtimev26.02Streaming speech-to-text
Voxtral TTSv26.03Zero-shot voice cloning

Batch or realtime

Choose by when you need the text. A finished video you want captioned is a batch job: upload the audio, wait, read the transcript. A live call or a stream you want subtitled as it plays needs a realtime model. For video production, batch is almost always the right choice, because the audio already exists in full and a batch model can see the whole file.

Prepare the audio before any transcription model

Whichever model you pick, cleaner input gives better text and smaller uploads. Sume's audio detach takes one Sume-hosted video and returns a new audio file for $0.01 per job. It defaults to sample-exact wav, accepts format of wav or mp3, and offers channels: "mono" and sample_rate of 16000, 44100 or 48000. The docs name 16000 Hz mono as the speech-to-text shape.

Source video can be up to 1800 seconds and the output up to 900 seconds, so a longer track needs a range. The video must already be on media.sume.com; import it first with POST /v1/media-imports.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: detach-voxtral-prep-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "channels": "mono",
    "sample_rate": 16000
  }'

If you just need a transcript on Sume

Video inspect can transcribe the clip's audio directly with transcribe: true. The documented transcript rate is $0.01 per audio minute, and segmentation.mode: "sentence" returns sentence segments shaped for caption lines. A silent clip fails with inspect_source_has_no_audio, so check probe.has_audio first.

Use an outside model like Voxtral when you want a specific vendor's transcription; check its own page for languages and limits.

Sources

Related posts

More in Models

All Models posts

Written by Sume