Dictation API: AssemblyAI cleaned text vs Sume STT word timings
AssemblyAI's Dictation API returns send-ready text. Sume STT returns a transcript with words[] timings and no cleanup flag, so you do the filler removal.

A dictation API is a speech-to-text call that returns text ready to send, not a raw transcript. AssemblyAI's Dictation API page says it returns text with filler gone, self-corrections resolved and names spelled right. Sume's STT 1.0 returns a transcript with words[] timings and has no cleanup option, so any filler or correction handling is a step you add after the job.
AssemblyAI facts are from its product page; Sume facts are from the API reference.
What does each service return?
AssemblyAI describes the Dictation API as the first API built for dictation: users speak and it returns finished text. Sume STT 1.0 completes with public-safe text, language fields when available, and words[] word-level timings as { word, start, end } in seconds from the audio start.
| Property | AssemblyAI Dictation API (product page) | Sume STT 1.0 (OpenAPI) |
|---|---|---|
| Output | Text ready to send | Transcript text plus words[] |
| Filler words | Gone | Not removed by a flag |
| Self-corrections | Resolved | Not resolved by a flag |
| Word timings | Not stated on the page | Always returned |
Can I turn on cleanup in Sume STT?
No. The request schema says word timings are always returned, that you send segmentation to also get sentence segments, and that provider knobs such as diarize and tag_audio_events are fixed server-side. There is no filler or self-correction option to set.
The request takes a public HTTPS audio_url, an optional language_code (omit it to auto-detect) and an optional duration_seconds from 1 to 600.
What do I do after the transcript comes back?
Filter in your own code. The timed words[] array lets you drop tokens you decide are filler and keep the rest in order. Resolving a self-correction ("Tuesday, no, Wednesday") is a language task and needs a rule or a text model on your side; the timings alone do not do it.
If your goal is a video without fillers rather than clean text, the work happens on the clip instead: see remove filler words from video. Captions are a separate surface; Video captions covers speech-to-captions, which needs audible speech.
When is raw plus timings the better fit?
When you need to know when each word was said: subtitles, highlight-as-you-read, or cutting audio at word boundaries. Cleaned text discards exactly that detail. When you only need a message to paste into a field, a dictation-style service matches the job more directly.
Sources
Related posts
More in Models
- AssemblyAI text to speech: coming soon, and what Sume has today
AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
- Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
Written by Sume