Scribe v2, Realtime or Medical: which model for interview clips

ElevenLabs lists Scribe v2 with up to 32 speakers and 90+ languages, Realtime at about 150 ms, and Medical with 35% fewer errors. Which fits interview clips.

4 min readSume
All posts

For interview clips, use Scribe v2 (batch): it handles up to 32 speakers across 90+ languages. Scribe v2 Realtime is for live use at about 150 ms, and Scribe v2 Medical is aimed at clinical audio. The ElevenLabs models page, read 2026-10-03, lists these three, and the vendor's 35 percent error claim applies to clinical audio only.

The three models

All figures come from the vendor page.

ElevenLabs Scribe models (read 2026-10-03)
ModelListed capabilities
Scribe v290+ languages, diarization up to 32 speakers, 65 entity types, keyterm prompting
Scribe v2 RealtimeAbout 150 ms
Scribe v2 MedicalClaims 35% fewer errors on clinical audio

Matching them to interview-to-clips

An interview to clips workflow transcribes a recording after the fact, finds quotable moments, and cuts them. That is a batch job, so Scribe v2 fits. Diarization matters because you want to attribute a quote to the right person; keyterm prompting helps with guest names and product terms.

Realtime is for captions during a live event. Medical is a specialist model; the 35 percent figure is the vendor's claim about clinical audio and says nothing about general interviews.

The Sume path

Sume has its own speech to text, Sume STT 1.0, at POST /v1/stt-1.0/transcribe, and the hosted MCP tool stt_create. Video inspect can transcribe one hosted clip at a public rate of $0.01 per audio minute, with an optional language_code hint and sentence segments[]. The docs reviewed here do not state a speaker count limit or diarization labels for Sume STT, so if you need speaker attribution, check the catalog first rather than assuming it.

Trim each chosen quote with the video trim route at a flat $0.02, and caption it with the captions route.

Keyterms and names

Interviews are full of names, companies and product terms. A keyterm list reduces the chance that a guest's name is misspelled in every caption. Prepare it before the call from the guest's bio and the topic list.

After transcription, spot-check the first mention of every name and the numbers, since an error repeats down every clip that quotes it.

A selection rule

Count speakers and languages first. If you need more than a few labelled speakers, confirm that your transcription path returns speaker labels before you build the cutting step on top of it. Keep the transcript and its model id with each clip so a quote can be traced to its source minute.

Sources

Related posts

More in Models

All Models posts

Written by Sume