AssemblyAI format_text false: 'one hundred' vs 100 and Sume STT

AssemblyAI's format_text: false returns spoken form, 'one hundred' not 100. Sume STT has no formatting toggle; its request is a closed schema.

4 min readSume
All posts

On August 20 AssemblyAI added spoken-form output for async universal-3-5-pro: set format_text: false and a speaker saying "one hundred" comes back as "one hundred", not "100". Sume STT has no format_text field or equivalent toggle, so you cannot choose between formatted and spoken form on the request.

AssemblyAI's behavior is from its changelog; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. The Sume docs do not say how numbers are written in the output.

What did AssemblyAI change?

The changelog says format_text: false "preserves the transcript in spoken form rather than formatted text: numbers, dates, and other entities are left unnormalized". It calls this useful for verbatim workflows such as medical dictation, and says it is available for all accounts.

What can I set on a Sume STT request?

audio_url, optional segmentation, language_code, duration_seconds, metadata and the job mode fields. The schema is closed, and provider knobs are fixed server-side. Nothing controls number or date normalization.

Spoken-form control, AssemblyAI changelog vs Sume STT schema, read 2026-10-01.
NeedAssemblyAI asyncSume `sume/stt-1.0`
Keep "one hundred" as spokenformat_text: false on universal-3-5-proNo field
Choose modelspeech_modelsPublic id sume/stt-1.0 only
Word timingsNot covered in this entrywords[], always returned
Sentence groupingNot covered in this entryOptional segmentation

Can I get verbatim text another way on Sume?

Only by testing. Run a clip that contains numbers, look at the words[] tokens, and see whether they are digits or words. The docs promise neither. If you need guaranteed spoken form, for example for medical dictation, pick a service that documents that switch.

What if I need digits from spoken words?

That is the easy direction: normalize in your own code from the tokens you receive. Each token keeps start and end, so the times survive your edit. The related Deepgram numerals post covers phone digits.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume