AssemblyAI disfluencies: true (English only) vs Sume STT words[]

AssemblyAI's disfluencies: true raised filler-word recall from 49.7% to 76.3%, English only. Sume STT has no such flag, so check words[] on your clip.

4 min readSume
All posts

AssemblyAI's August 27 changelog says async disfluencies: true on Universal-3.5 Pro now captures filler words at parity with Universal-2, with recall up from 49.7% to 76.3% in internal benchmarks, English only. Sume STT has no disfluency flag and its docs do not say whether fillers are kept, so measure it on your own audio.

AssemblyAI's figures are from its changelog; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. Both numbers in the changelog are AssemblyAI's own internal benchmark.

What exactly changed?

The entry says the update uses the same mechanism as Universal-2 and requires no change beyond setting disfluencies: true. It states support is currently English-only. A separate dictation guide covers the opposite, filler-free text; see the dictation post.

What does Sume expose?

Word timings are always returned, with no flag, as words[] entries of { word, start, end } in seconds. Provider knobs are fixed server-side, and a disfluency switch is not among the request fields.

Filler-word handling, AssemblyAI changelog vs Sume STT schema, read 2026-10-01.
ItemAssemblyAI asyncSume `sume/stt-1.0`
Switch for fillersdisfluencies: trueNone
LanguagesEnglish only (currently)Not published; hint via language_code
Recall figure49.7% to 76.3%, internal benchmarkNone published
Where fillers show upIn the transcriptDocs do not say

How do I check fillers on Sume?

Send a short clip you have marked up by hand, with a few known "um" and "uh" words, then search words[] for them. The array gives the start and end time of each token, so you can also measure pauses. Do not assume fillers are kept or dropped; the docs are silent.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-demo-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
    "duration_seconds": 120,
    "language_code": "en"
  }'

Why would anyone want fillers?

Editing and speech coaching use them: a cut list built from timestamps can trim hesitations, and speech-rate math counts them or not, depending on what you want to measure. If a verbatim record is a requirement, a documented flag is safer than a measured behavior.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume