Scribe v2 audio-event tags and 32 speakers vs Sume STT's fixed flags

ElevenLabs Scribe v2 tags audio events and separates up to 32 speakers. Sume STT fixes diarize and tag_audio_events server-side, so those toggles do not exist.

5 min readSume
All posts

ElevenLabs Scribe v2 supports audio-event tagging (laughter, applause and similar) and speaker identification for up to 32 speakers, and Sume STT gives you neither as a switch: the OpenAPI contract says provider knobs such as diarize and tag_audio_events are fixed server-side. What you do get from Sume is the transcript text, a detected language with probability when available, and word-level timings.

The Scribe facts are from ElevenLabs' speech-to-text documentation (read 2026-10-10). The Sume facts are from the Sume OpenAPI contract.

What the two pages promise

On the ElevenLabs page, Scribe v2, Scribe v2 Realtime and Scribe v2 Medical are listed; the page claims 90+ languages, a 3 GB and 10-hour file ceiling (1 hour for multichannel), word-level timestamps, keyterm prompting of up to 1,000 terms in batch, and about 150 ms latency for the realtime model. Sume's STT request has audio_url, language_code, duration_seconds, segmentation, metadata and the delivery fields, and that is all.

Speaker and event features (ElevenLabs read 2026-10-10; Sume per OpenAPI)
FeatureElevenLabs Scribe v2Sume STT 1.0
Speaker identificationUp to 32 speakersNo request switch; diarize is fixed server-side
Audio event tagsSupportedNo request switch; tag_audio_events is fixed server-side
Word timingsYesAlways returned as words[] with start and end seconds
Sentence segmentsNot statedOptional segmentation.mode: sentence, gapless
Max job length10 hoursduration_seconds 1 to 600

Working with fixed flags

Fixed flags mean the word list can contain provider token types. The OpenAPI schema says each word entry may carry a type, for example word or spacing, when the provider supplies it. Filter on type if you build captions, so spacing tokens do not become subtitle words.

The contract does not promise event or speaker labels in the output, so do not build a feature on them. If you need speaker separation, the Sume-compatible workaround is structural: record or export one track per speaker and submit one job per track at $0.01 per audio minute.

Per-track speaker jobs in practice

Take a two-person podcast with separate tracks of 30 minutes each. Each track is 1,800 seconds, so it needs three STT windows of 600 seconds, six jobs in all, at about 60 cents of STT if every window is a full 10 minutes ($0.01 per minute times 60 minutes). You then merge the two word lists by start time to get a speaker-attributed transcript, because you know which track each word came from.

That is more work than a diarize flag, but it is deterministic: the speaker label is a fact of your file layout rather than a model guess.

Which one to use

Use Scribe when you have a single mixed file with many speakers and want tags. Use Sume STT when you want transcripts as one step in a chain with detach, captions or TTS, and you can live with fixed settings. Sume's captions job reuses STT word timings and, per its docs, runs speech-to-text only when you do not send your own words.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume