AssemblyAI disfluencies: true (English only) vs Sume STT words[]
AssemblyAI's disfluencies: true raised filler-word recall from 49.7% to 76.3%, English only. Sume STT has no such flag, so check words[] on your clip.

AssemblyAI's August 27 changelog says async disfluencies: true on Universal-3.5 Pro now captures filler words at parity with Universal-2, with recall up from 49.7% to 76.3% in internal benchmarks, English only. Sume STT has no disfluency flag and its docs do not say whether fillers are kept, so measure it on your own audio.
AssemblyAI's figures are from its changelog; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. Both numbers in the changelog are AssemblyAI's own internal benchmark.
What exactly changed?
The entry says the update uses the same mechanism as Universal-2 and requires no change beyond setting disfluencies: true. It states support is currently English-only. A separate dictation guide covers the opposite, filler-free text; see the dictation post.
What does Sume expose?
Word timings are always returned, with no flag, as words[] entries of { word, start, end } in seconds. Provider knobs are fixed server-side, and a disfluency switch is not among the request fields.
| Item | AssemblyAI async | Sume `sume/stt-1.0` |
|---|---|---|
| Switch for fillers | disfluencies: true | None |
| Languages | English only (currently) | Not published; hint via language_code |
| Recall figure | 49.7% to 76.3%, internal benchmark | None published |
| Where fillers show up | In the transcript | Docs do not say |
How do I check fillers on Sume?
Send a short clip you have marked up by hand, with a few known "um" and "uh" words, then search words[] for them. The array gives the start and end time of each token, so you can also measure pauses. Do not assume fillers are kept or dropped; the docs are silent.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"language_code": "en"
}'Why would anyone want fillers?
Editing and speech coaching use them: a cut list built from timestamps can trim hesitations, and speech-rate math counts them or not, depending on what you want to measure. If a verbatim record is a requirement, a documented flag is safer than a measured behavior.
Sources
Related posts
More in Developers
- AssemblyAI format_text false: 'one hundred' vs 100 and Sume STT
AssemblyAI's format_text: false returns spoken form, 'one hundred' not 100. Sume STT has no formatting toggle; its request is a closed schema.
- AssemblyAI speech models: unpinned routing and Sume's fixed STT id
AssemblyAI will route unpinned async requests to Universal-3.5 Pro; pin speech_models to stay put. What changes, and why Sume STT has no model to pin.
- AssemblyAI Sync API: one call, and Sume STT sync mode
AssemblyAI's Sync API returns a short-clip transcript in one POST. Sume STT mode sync waits up to 30 seconds, then returns the job id to poll if not done.
- AssemblyAI Sync API file limit vs Sume STT duration_seconds
AssemblyAI's launch post frames its Sync API for short clips. Sume STT 1.0 takes duration_seconds from 1 to 600 and reserves one minute if you omit it.
Written by Sume