Gladia diarization vs Sume speech-to-text: what you can switch
Gladia lists speaker diarization and word timestamps on every plan. Sume STT always returns word timings but fixes diarization server-side, with no switch.

Sume speech-to-text does not give you a diarization switch the way Gladia's feature list does. Sume STT 1.0 always returns word timings, optionally adds sentence segments, and fixes provider options such as diarize and tag_audio_events on the server, so sending either field is refused with a 400.
Gladia's own pricing page (read 2026-10-10) lists speaker diarization, word-level timestamps, automatic language detection and 100+ languages on all plans. This post sets those lines next to what the Sume request body accepts, so you can decide before you wire either one into a pipeline. Sume facts come from the API reference and the STT request schema in the API source.
What each request lets you control
The useful comparison is the set of knobs, not the headline price. Gladia describes a feature set that applies to every plan. Sume describes a request body with a short list of fields and a stated list of things it will not let you send.
| Question | Gladia (read 2026-10-10) | Sume STT 1.0 |
|---|---|---|
| Word-level timestamps | Listed as a core feature on all plans | Always returned; there is no flag to enable them |
| Speaker diarization | Listed as a core feature on all plans | Fixed server-side; diarize is refused if you send it |
| Language | Automatic language detection, 100+ languages | Optional language_code; omit it to auto-detect |
| Sentence segments | Not listed on the pricing page | Opt in with segmentation: { mode: "sentence" } |
| Input | Not stated on the pricing page | audio_url, a public HTTPS URL |
| Concurrency | 25 async requests on Starter, more on Growth | See rate limits in the Sume docs |
What Sume STT takes and returns
The Sume body is deliberately small: audio_url, an optional language_code, an optional duration_seconds between 1 and 600, an optional segmentation block, plus the usual mode, webhook_url and wait_timeout_seconds fields. The schema is strict, so any other field is a 400 rather than a silent no-op.
duration_seconds is only a reservation hint. If you omit it, Sume reserves one minute of usage; if you send it, the reservation matches your clip. The API description also says segmentation adds sentence segments on top of the word timings that always come back.
If your source is a video, audio detach is the documented step before STT. It returns a wav by default, and the docs call 16000 Hz with channels: "mono" the STT shape. Detach costs $0.01 per job and caps the source at 1800 seconds, so a long recording needs a range.
- No
diarizeortag_audio_eventsfield: both are fixed server-side. - No per-request language list; one optional
language_codehint. - Results arrive as a job, so read them through Jobs and results or a webhook.
Price for the same hour
Both vendors land in a similar range for ordinary audio, so price rarely decides this one. Gladia's Starter plan is $0.61 per hour for async transcription and $0.75 per hour real-time, and its page says committed Growth pricing starts as low as $0.20 per hour async. Sume STT lists at $0.01 per audio minute, which is $0.60 for 60 minutes. The golden pricing test bills a 61 second clip at 1.0167 audio minutes, or $0.010168, so partial minutes are prorated rather than rounded up.
That makes a 10-hour archive $6.00 on Sume at list and $6.10 on Gladia Starter. The $0.10 gap is not the deciding factor. What matters is whether you need diarization you can control, and Gladia states it as a feature while Sume does not expose it as a choice.
Which one fits which job
Pick Gladia when the transcript itself is the product and you want a vendor that advertises diarization, 100+ languages and compliance claims on the plan page. Its page also lists GDPR, HIPAA and SOC 2 Type 2, which I did not find in the Sume docs read for this post.
Pick Sume STT when the transcript feeds other Sume steps. The same media pipeline gives you video captions, where speech-to-text timing is the source of truth for burned-in words, plus trim, timeline and webhooks under one key. A clip with no speech fails with caption_no_speech rather than returning an empty caption.
A short decision list
Before choosing, write down the three things your output needs.
- Need speaker labels you can turn on or off: use Gladia, per its page.
- Need word timings and sentence segments feeding captions or cuts inside Sume: use STT 1.0.
- Need a language outside what you have verified on each side: test one clip first, because the Sume docs read this session do not publish a language count.
Sources
Related posts
More in Comparisons
- Google Vids point-in-time editing: timed text via Sume cues
Google Vids now shows only what is on screen at the playhead. Sume has no canvas, but caption cues with start and end place timed text by API.
- Grok Imagine's 4 keyframes and 7 references vs Sume's one image
xAI's Grok Imagine 1.5 takes up to 4 keyframes and up to 7 references. Sume's grok-imagine-video-1.5 row takes one image only. Rows to use for multi-image work.
- Grok Imagine Lite upscales to 1080p: what Sume offers instead
xAI describes Grok Imagine Video 1.5 Lite as lowest cost with upscaled 1080p. Sume has no Lite row; it lists grok-imagine-video-1.5 at a flat rate. Compare.
- Grok Imagine's 3 voice references vs Sume reference-audio rows
xAI's Grok Imagine 1.5 takes up to 3 voice references. Sume's Grok row takes none; Seedance, Wan 3.0 and MiniMax accept reference audio under limits.
Written by Sume