Scribe v2, Realtime or Medical: which model for interview clips
ElevenLabs lists Scribe v2 with up to 32 speakers and 90+ languages, Realtime at about 150 ms, and Medical with 35% fewer errors. Which fits interview clips.

For interview clips, use Scribe v2 (batch): it handles up to 32 speakers across 90+ languages. Scribe v2 Realtime is for live use at about 150 ms, and Scribe v2 Medical is aimed at clinical audio. The ElevenLabs models page, read 2026-10-03, lists these three, and the vendor's 35 percent error claim applies to clinical audio only.
The three models
All figures come from the vendor page.
| Model | Listed capabilities |
|---|---|
| Scribe v2 | 90+ languages, diarization up to 32 speakers, 65 entity types, keyterm prompting |
| Scribe v2 Realtime | About 150 ms |
| Scribe v2 Medical | Claims 35% fewer errors on clinical audio |
Matching them to interview-to-clips
An interview to clips workflow transcribes a recording after the fact, finds quotable moments, and cuts them. That is a batch job, so Scribe v2 fits. Diarization matters because you want to attribute a quote to the right person; keyterm prompting helps with guest names and product terms.
Realtime is for captions during a live event. Medical is a specialist model; the 35 percent figure is the vendor's claim about clinical audio and says nothing about general interviews.
The Sume path
Sume has its own speech to text, Sume STT 1.0, at POST /v1/stt-1.0/transcribe, and the hosted MCP tool stt_create. Video inspect can transcribe one hosted clip at a public rate of $0.01 per audio minute, with an optional language_code hint and sentence segments[]. The docs reviewed here do not state a speaker count limit or diarization labels for Sume STT, so if you need speaker attribution, check the catalog first rather than assuming it.
Trim each chosen quote with the video trim route at a flat $0.02, and caption it with the captions route.
Keyterms and names
Interviews are full of names, companies and product terms. A keyterm list reduces the chance that a guest's name is misspelled in every caption. Prepare it before the call from the guest's bio and the topic list.
After transcription, spot-check the first mention of every name and the numbers, since an error repeats down every clip that quotes it.
A selection rule
Count speakers and languages first. If you need more than a few labelled speakers, confirm that your transcription path returns speaker labels before you build the cutting step on top of it. Keep the transcript and its model id with each clip so a quote can be traced to its source minute.
Sources
Related posts
More in Models
- Seedance 2.5 on BytePlus excludes the US: check access first
BytePlus lists Seedance 2.5 on ModelArk in supported markets excluding the United States. What that means for a US team, and how to read the Sume catalog.
- Self-host MiniMax H3 or use an API: a decision table
MiniMax released H3 as an open-weight omni-modal model on Jul 31, 2026. A decision table for running it yourself versus calling a hosted video API.
- Veo 3.1 extend adds 7 s up to 20 times, 720p only: plan a long clip
Google's Veo docs let you extend a clip by 7 seconds up to 20 times, at 720p only. Read the arithmetic, then compare it with one-call 30 second models on Sume.
- Voxtral Mini Transcribe 2 and Realtime v26.02: what Mistral lists
Mistral lists Voxtral Mini Transcribe 2, Voxtral Realtime v26.02 and Voxtral TTS v26.03. How to prepare video audio for any transcription model.
Written by Sume