Gladia diarization vs Sume speech-to-text: what you can switch

Gladia lists speaker diarization and word timestamps on every plan. Sume STT always returns word timings but fixes diarization server-side, with no switch.

5 min readSume
All posts

Sume speech-to-text does not give you a diarization switch the way Gladia's feature list does. Sume STT 1.0 always returns word timings, optionally adds sentence segments, and fixes provider options such as diarize and tag_audio_events on the server, so sending either field is refused with a 400.

Gladia's own pricing page (read 2026-10-10) lists speaker diarization, word-level timestamps, automatic language detection and 100+ languages on all plans. This post sets those lines next to what the Sume request body accepts, so you can decide before you wire either one into a pipeline. Sume facts come from the API reference and the STT request schema in the API source.

What each request lets you control

The useful comparison is the set of knobs, not the headline price. Gladia describes a feature set that applies to every plan. Sume describes a request body with a short list of fields and a stated list of things it will not let you send.

Gladia features from its pricing page (read 2026-10-10); Sume from the STT 1.0 request schema and OpenAPI description
QuestionGladia (read 2026-10-10)Sume STT 1.0
Word-level timestampsListed as a core feature on all plansAlways returned; there is no flag to enable them
Speaker diarizationListed as a core feature on all plansFixed server-side; diarize is refused if you send it
LanguageAutomatic language detection, 100+ languagesOptional language_code; omit it to auto-detect
Sentence segmentsNot listed on the pricing pageOpt in with segmentation: { mode: "sentence" }
InputNot stated on the pricing pageaudio_url, a public HTTPS URL
Concurrency25 async requests on Starter, more on GrowthSee rate limits in the Sume docs

What Sume STT takes and returns

The Sume body is deliberately small: audio_url, an optional language_code, an optional duration_seconds between 1 and 600, an optional segmentation block, plus the usual mode, webhook_url and wait_timeout_seconds fields. The schema is strict, so any other field is a 400 rather than a silent no-op.

duration_seconds is only a reservation hint. If you omit it, Sume reserves one minute of usage; if you send it, the reservation matches your clip. The API description also says segmentation adds sentence segments on top of the word timings that always come back.

If your source is a video, audio detach is the documented step before STT. It returns a wav by default, and the docs call 16000 Hz with channels: "mono" the STT shape. Detach costs $0.01 per job and caps the source at 1800 seconds, so a long recording needs a range.

  • No diarize or tag_audio_events field: both are fixed server-side.
  • No per-request language list; one optional language_code hint.
  • Results arrive as a job, so read them through Jobs and results or a webhook.

Price for the same hour

Both vendors land in a similar range for ordinary audio, so price rarely decides this one. Gladia's Starter plan is $0.61 per hour for async transcription and $0.75 per hour real-time, and its page says committed Growth pricing starts as low as $0.20 per hour async. Sume STT lists at $0.01 per audio minute, which is $0.60 for 60 minutes. The golden pricing test bills a 61 second clip at 1.0167 audio minutes, or $0.010168, so partial minutes are prorated rather than rounded up.

That makes a 10-hour archive $6.00 on Sume at list and $6.10 on Gladia Starter. The $0.10 gap is not the deciding factor. What matters is whether you need diarization you can control, and Gladia states it as a feature while Sume does not expose it as a choice.

Which one fits which job

Pick Gladia when the transcript itself is the product and you want a vendor that advertises diarization, 100+ languages and compliance claims on the plan page. Its page also lists GDPR, HIPAA and SOC 2 Type 2, which I did not find in the Sume docs read for this post.

Pick Sume STT when the transcript feeds other Sume steps. The same media pipeline gives you video captions, where speech-to-text timing is the source of truth for burned-in words, plus trim, timeline and webhooks under one key. A clip with no speech fails with caption_no_speech rather than returning an empty caption.

A short decision list

Before choosing, write down the three things your output needs.

  • Need speaker labels you can turn on or off: use Gladia, per its page.
  • Need word timings and sentence segments feeding captions or cuts inside Sume: use STT 1.0.
  • Need a language outside what you have verified on each side: test one clip first, because the Sume docs read this session do not publish a language count.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume