Medical transcription API: Scribe v2 Medical GA, and what Sume STT is
ElevenLabs reports Scribe v2 Medical as GA on 2026-09-11. Sume's STT is general speech-to-text with word timings, not a medical product. Where that matters.

If you need transcription tuned for clinical speech, Sume's speech-to-text is not that product. A changelog roundup reports that ElevenLabs made Scribe v2 Medical generally available on 2026-09-11. Sume STT is a general transcription job: you give it a public HTTPS audio URL and get back text, word timings and an optional language hint. Its docs make no medical, clinical-vocabulary or health-data compliance claim, so do not assume one.
Use this page to decide quickly whether Sume STT can serve your case at all. Sume also does not publish accuracy figures for clinical speech in the docs read here, so there is nothing to compare against ElevenLabs' medical model.
What does Sume STT do?
A job takes audio_url, an optional language_code (omit it for auto-detect), and an optional duration_seconds between 1 and 600 that improves the usage reservation. Word timings are always returned, and segmentation with sentence mode adds sentence segments. Provider settings such as diarization and audio-event tagging are fixed on the server, so you cannot switch speaker labels on or add a custom term list through the request.
| Need | Sume STT | Where to look |
|---|---|---|
| Clinical vocabulary tuning | Not documented | A medical-specific vendor |
| Health-data compliance statements | Not made in the docs | Sume's terms, and your counsel |
| Speaker labels | Not settable in the request | A tool with diarization controls |
| Word timings | Always returned | Result words[] |
| Audio length per job | Up to 600 seconds reserved | Split longer audio |
When is general STT still fine?
For non-clinical speech near a clinic: a marketing video with a doctor, a patient-education explainer, an internal training recording, or captions for a conference talk. Check every transcript yourself before publishing, because a general model may misspell drug names and anatomy. Do not feed patient recordings into any service until your privacy review says you may. A useful test is to ask whether a wrong word would change a decision; if it would, a human must review the transcript, whichever tool made it.
What should I do?
Write down the one requirement you cannot waive: a vocabulary, a compliance statement or speaker labels. If any of them applies, shortlist a medical-specific tool and ask for its terms. If none applies, run a ten-minute sample through Sume STT and proofread it. The API reference lists the STT request fields, and Jobs and results shows how to read the result.
Sources
Related posts
More in Comparisons
- Seedance 2.0 4K on Hedra: what Sume's seedance-2 lists
Hedra lists Seedance 2.0 at 4K for about $9.07 a minute. Sume's docs list seedance-2 at 480p to 1080p; here is where 4K does and does not work on Sume.
- Segmind API vs Sume: PixelFlow workflows or saved Formats
Segmind turns visual PixelFlow graphs into API endpoints. Sume saves an agent thread as a Format you call over the API. How the two reuse recipes.
- Sonilo segment-level music controls vs Sume section markers
Sonilo's text-to-music lets you set styles and moods per section. On Sume you write section markers like [0:00-0:30] Intro: inside one 5000-character prompt.
- Sonilo video-to-music on fal.ai: 600 s of footage vs Sume's route
Sonilo's video-to-music model scores footage up to 600 seconds on fal.ai. Sume has no video-to-music call: inspect the clip, write a prompt, mix with Timeline.
Written by Sume