Deepgram Flux numerals toggle vs Sume STT digits
Deepgram Flux can switch numerals on mid-stream for PINs and phone numbers. Sume STT is a batch job with fixed provider settings and no numerals flag.

Sume STT has no numerals setting and no stream to toggle it on. It is a batch job: you submit an audio_url, and the transcript comes back with whatever formatting the fixed provider settings produce. Deepgram's Flux STT, by contrast, lets you send a Configure message mid-stream to switch to digits for a PIN or phone number.
Deepgram facts are from its changelog; Sume facts are from the API reference. Both read 2026-10-01.
What did Deepgram add to Flux?
The Sep 25, 2026 changelog entry says you can switch Flux STT to digits for a PIN, phone number or order number, then back to words, without reconnecting. You send { "type": "Configure", "numerals": true }. The update applies to transcripts Flux sends after it processes the message, and the numerals query parameter still sets the initial value.
How does Sume STT handle this?
Sume has no equivalent control. The schema states that provider knobs such as diarize and tag_audio_events are fixed server-side, and the only language control is an optional BCP-47 hint. With mode: async the call returns immediately with status_url, so there is no open connection to reconfigure.
| Aspect | Deepgram Flux | Sume STT 1.0 |
|---|---|---|
| Digits toggle | numerals boolean, changeable mid-stream | No flag |
| Transport | Streaming connection | Job with status_url and result_url |
| Word timings | Not stated in the entry | Always returned |
How do I find a spoken PIN or phone number in a Sume transcript?
Read words[]. Each entry has word, start and end, so you can locate the digits in the audio. Then normalize in your own code: map spelled-out digits to numerals, or the reverse, after the job completes. Do not assume the transcript will already be in digit form.
What should I do next?
If digits must be correct live, in a call flow, use a streaming service that offers the toggle. For recorded audio, normalize after transcription and verify against word timings. See also real-time speech-to-text on Sume.
Sources
Related posts
More in Developers
- Deepgram Flux TTS speed 0.5 to 1.5 vs Sume's 0.6 to 1.5
Deepgram widened Flux TTS speed to 0.5-1.5 in 0.05 steps on 2026-08-31. Sume TTS accepts generation_config.speed from 0.6 to 1.5, plus a deprecated enum.
- Devin Desktop ACP session/new mcp_servers with Sume's remote URL
Devin Desktop 3.10.35 uses MCP servers an ACP client passes in session/new. Pass Sume's streamable HTTP URL and an API-key header, not editor config.
- ElevenLabs API timeout: cascade_timeout_seconds vs Sume waits
ElevenLabs added cascade_timeout_seconds (2-15 s) for Speech Engine retries. Sume's wait_timeout_seconds is a different knob: a 0-30 s HTTP wait on a job.
- ElevenLabs API cursor pagination, and how Sume pages lists
ElevenLabs added a cursor-paginated phone number endpoint. Sume list routes use keyset pages: pass next_cursor back as cursor until has_more is false.
Written by Sume