Scribe v2 audio-event tags and 32 speakers vs Sume STT's fixed flags
ElevenLabs Scribe v2 tags audio events and separates up to 32 speakers. Sume STT fixes diarize and tag_audio_events server-side, so those toggles do not exist.

ElevenLabs Scribe v2 supports audio-event tagging (laughter, applause and similar) and speaker identification for up to 32 speakers, and Sume STT gives you neither as a switch: the OpenAPI contract says provider knobs such as diarize and tag_audio_events are fixed server-side. What you do get from Sume is the transcript text, a detected language with probability when available, and word-level timings.
The Scribe facts are from ElevenLabs' speech-to-text documentation (read 2026-10-10). The Sume facts are from the Sume OpenAPI contract.
What the two pages promise
On the ElevenLabs page, Scribe v2, Scribe v2 Realtime and Scribe v2 Medical are listed; the page claims 90+ languages, a 3 GB and 10-hour file ceiling (1 hour for multichannel), word-level timestamps, keyterm prompting of up to 1,000 terms in batch, and about 150 ms latency for the realtime model. Sume's STT request has audio_url, language_code, duration_seconds, segmentation, metadata and the delivery fields, and that is all.
| Feature | ElevenLabs Scribe v2 | Sume STT 1.0 |
|---|---|---|
| Speaker identification | Up to 32 speakers | No request switch; diarize is fixed server-side |
| Audio event tags | Supported | No request switch; tag_audio_events is fixed server-side |
| Word timings | Yes | Always returned as words[] with start and end seconds |
| Sentence segments | Not stated | Optional segmentation.mode: sentence, gapless |
| Max job length | 10 hours | duration_seconds 1 to 600 |
Working with fixed flags
Fixed flags mean the word list can contain provider token types. The OpenAPI schema says each word entry may carry a type, for example word or spacing, when the provider supplies it. Filter on type if you build captions, so spacing tokens do not become subtitle words.
The contract does not promise event or speaker labels in the output, so do not build a feature on them. If you need speaker separation, the Sume-compatible workaround is structural: record or export one track per speaker and submit one job per track at $0.01 per audio minute.
Per-track speaker jobs in practice
Take a two-person podcast with separate tracks of 30 minutes each. Each track is 1,800 seconds, so it needs three STT windows of 600 seconds, six jobs in all, at about 60 cents of STT if every window is a full 10 minutes ($0.01 per minute times 60 minutes). You then merge the two word lists by start time to get a speaker-attributed transcript, because you know which track each word came from.
That is more work than a diarize flag, but it is deterministic: the speaker label is a fact of your file layout rather than a model guess.
Which one to use
Use Scribe when you have a single mixed file with many speakers and want tags. Use Sume STT when you want transcripts as one step in a chain with detach, captions or TTS, and you can live with fixed settings. Sume's captions job reuses STT word timings and, per its docs, runs speech-to-text only when you do not send your own words.
Sources
Related posts
More in Comparisons
- ElevenLabs sound effects: prompt influence, 48 kHz WAV, and Sume
ElevenLabs SFX offers high or low prompt influence and 48 kHz WAV. Sume has no sound-effects route; a Music prompt plus a split makes a sting. What you lose.
- Gemini TTS has 30 prebuilt voices; Sume TTS takes avatar voice ids
Google's Gemini TTS page lists 30 prebuilt voices. Sume TTS takes a voice UUID or voi_ id, or an avatar whose voice.status is ready; cloning stays app-only.
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
- Grok Imagine Lite upscales to 1080p: what Sume offers instead
xAI describes Grok Imagine Video 1.5 Lite as lowest cost with upscaled 1080p. Sume has no Lite row; it lists grok-imagine-video-1.5 at a flat rate. Compare.
Written by Sume