AssemblyAI format_text false: 'one hundred' vs 100 and Sume STT
AssemblyAI's format_text: false returns spoken form, 'one hundred' not 100. Sume STT has no formatting toggle; its request is a closed schema.

On August 20 AssemblyAI added spoken-form output for async universal-3-5-pro: set format_text: false and a speaker saying "one hundred" comes back as "one hundred", not "100". Sume STT has no format_text field or equivalent toggle, so you cannot choose between formatted and spoken form on the request.
AssemblyAI's behavior is from its changelog; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. The Sume docs do not say how numbers are written in the output.
What did AssemblyAI change?
The changelog says format_text: false "preserves the transcript in spoken form rather than formatted text: numbers, dates, and other entities are left unnormalized". It calls this useful for verbatim workflows such as medical dictation, and says it is available for all accounts.
What can I set on a Sume STT request?
audio_url, optional segmentation, language_code, duration_seconds, metadata and the job mode fields. The schema is closed, and provider knobs are fixed server-side. Nothing controls number or date normalization.
| Need | AssemblyAI async | Sume `sume/stt-1.0` |
|---|---|---|
| Keep "one hundred" as spoken | format_text: false on universal-3-5-pro | No field |
| Choose model | speech_models | Public id sume/stt-1.0 only |
| Word timings | Not covered in this entry | words[], always returned |
| Sentence grouping | Not covered in this entry | Optional segmentation |
Can I get verbatim text another way on Sume?
Only by testing. Run a clip that contains numbers, look at the words[] tokens, and see whether they are digits or words. The docs promise neither. If you need guaranteed spoken form, for example for medical dictation, pick a service that documents that switch.
What if I need digits from spoken words?
That is the easy direction: normalize in your own code from the tokens you receive. Each token keeps start and end, so the times survive your edit. The related Deepgram numerals post covers phone digits.
Sources
Related posts
More in Developers
- AssemblyAI speech models: unpinned routing and Sume's fixed STT id
AssemblyAI will route unpinned async requests to Universal-3.5 Pro; pin speech_models to stay put. What changes, and why Sume STT has no model to pin.
- AssemblyAI Sync API: one call, and Sume STT sync mode
AssemblyAI's Sync API returns a short-clip transcript in one POST. Sume STT mode sync waits up to 30 seconds, then returns the job id to poll if not done.
- AssemblyAI Sync API file limit vs Sume STT duration_seconds
AssemblyAI's launch post frames its Sync API for short clips. Sume STT 1.0 takes duration_seconds from 1 to 600 and reserves one minute if you omit it.
- AssemblyAI Universal-3.5 Pro and Urdu: language_code on Sume STT
AssemblyAI fixed a bug that sent Urdu to Universal-3.5 Pro, which covers 18 languages. How a language hint behaves on Sume STT, and what to check.
Written by Sume