Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.

Mistral's model overview lists Voxtral Mini Transcribe 2 and Voxtral Realtime, both versioned v26.02, alongside Voxtral TTS v26.03 with zero-shot voice cloning. The page is the place to start for speech-to-text with Mistral, but it gives little else about limits, so confirm the details on each model's own page before you build.
What the overview lists
The overview groups the audio models by task. The transcription entries cover a batch model and a realtime model, and the speech entry covers synthesis with zero-shot voice cloning.
| Model | Version | Listed capability |
|---|---|---|
| Voxtral Mini Transcribe 2 | v26.02 | Speech-to-text |
| Voxtral Realtime | v26.02 | Streaming speech-to-text |
| Voxtral TTS | v26.03 | Zero-shot voice cloning |
Batch or realtime
Choose by when you need the text. A finished video you want captioned is a batch job: upload the audio, wait, read the transcript. A live call or a stream you want subtitled as it plays needs a realtime model. For video production, batch is almost always the right choice, because the audio already exists in full and a batch model can see the whole file.
Prepare the audio before any transcription model
Whichever model you pick, cleaner input gives better text and smaller uploads. Sume's audio detach takes one Sume-hosted video and returns a new audio file for $0.01 per job. It defaults to sample-exact wav, accepts format of wav or mp3, and offers channels: "mono" and sample_rate of 16000, 44100 or 48000. The docs name 16000 Hz mono as the speech-to-text shape.
Source video can be up to 1800 seconds and the output up to 900 seconds, so a longer track needs a range. The video must already be on media.sume.com; import it first with POST /v1/media-imports.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: detach-voxtral-prep-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"channels": "mono",
"sample_rate": 16000
}'If you just need a transcript on Sume
Video inspect can transcribe the clip's audio directly with transcribe: true. The documented transcript rate is $0.01 per audio minute, and segmentation.mode: "sentence" returns sentence segments shaped for caption lines. A silent clip fails with inspect_source_has_no_audio, so check probe.has_audio first.
Use an outside model like Voxtral when you want a specific vendor's transcription; check its own page for languages and limits.
Sources
Related posts
More in Models
- Wan 3.0 doubles clip length to 30 seconds: Alibaba ids vs Sume
Alibaba Model Studio lists wan3.0-video and wan3.0-video-prime at 2 to 30 seconds, up from 15 on Wan 2.7. What Sume's wan-3.0 accepts and a 30 s request.
- What is FLUX 3? The family map: image, video, audio, action
BFL's docs describe FLUX 3 as one family covering image, video with synchronized audio, audio and action. Which pieces have open weights and which do not.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume