Azure Speech SDK 1.52 adds streaming model choice: batch STT instead?
Azure's Speech SDK 1.52 lets you pick a streaming model for transcription and adds PushAudioInputStream commit in preview. When recorded-file jobs are enough.

Azure's Speech SDK 1.52 adds support for selecting a streaming model for speech transcription and a preview of commit support for PushAudioInputStream. Both are for live audio, so if your audio already exists as a file, a batch job with word timings is simpler and Sume's STT route is one option, with no streaming at all.
Azure details come from its Speech SDK release notes, read 2026-10-03, and Sume details from the OpenAPI reference.
What is in Speech SDK 1.52?
The notes list, under new features, commit support for PushAudioInputStream (marked preview) and support for selecting a streaming model for speech transcription. The JavaScript SDK 1.52 also gains the streaming-model selection. Bug fixes include a false embedded TTS timeout when the system clock changes during synthesis, a TTS crash with compressed audio when GStreamer is not found, and, on Android, a TTS crash when the codec returns a null buffer. The JavaScript SDK also fixes audio offset calculation for reliable reconnect and missing channel information in ConversationTranscriber results.
The notes do not name the selectable models or give prices, so check the Azure Speech pages for those before planning a migration.
| Item | Detail from the release notes |
|---|---|
| Streaming model selection | New in 1.52, also in the JavaScript SDK |
| PushAudioInputStream commit | New, preview |
| Reconnect | JavaScript fix for audio offset calculation |
| ConversationTranscriber | JavaScript fix for missing channel information |
| TTS fixes | Embedded timeout on clock change; compressed-audio crash without GStreamer |
Do you need streaming at all?
Streaming matters when someone is waiting on partial text: live captions, a voice agent, a dictation box. It does not matter for a podcast, a recorded meeting or a folder of customer calls. For those, you pay for a socket, reconnect logic and an audio push loop that a job on a finished file does not need.
A quick test: does anything you build change its behavior before the recording ends? If not, batch is enough.
What does the Sume alternative look like?
Sume STT is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and a few optional fields: language_code, duration_seconds from 1 to 600, sentence segmentation, metadata and the mode fields. It returns text and word-level words[] timings in seconds. A sync wait lasts at most 30 seconds, so for anything longer you submit async and poll status_url as described in the jobs docs. The public price is $0.01 per audio minute, and one minute is reserved if you leave out duration_seconds.
There is no push stream, no partial results and no model selector in the request. The reference says provider knobs are fixed server-side. If you need to pick a streaming model or push audio as it arrives, Azure's SDK is the right tool.
How do you choose?
Decide on the shape of the audio first, then the model. Live and interactive audio belongs on a streaming API. Finished files belong on batch jobs, where the unit is one job with a record.
- Live captions or voice agent: Azure Speech SDK streaming, with the new model choice.
- Recorded files up to ten minutes each: Sume STT jobs; split longer ones first.
- Both needs: stream live for the interface, then re-transcribe the final recording as a job for the archive.
- Always keep the audio and the transcript linked by an id you control.
Sources
Related posts
More in Comparisons
- Claude Code mod vs MCP server vs skill vs hook: where Sume fits
Claude Code mods, MCP servers, skills and settings hooks overlap. Which one gives an agent Sume's image and video tools, and which one only guards them.
- Comfy Agent Ask or Auto mode vs unattended Sume Format runs
Comfy Agent asks before each run or runs on its own. Sume API runs never ask. See which controls replace the approval prompt when a batch goes overnight.
- Comfy Agent vs an MCP agent for image and video work
Comfy Agent builds and runs ComfyUI graphs for you in Comfy Cloud; an MCP agent calls hosted tools like Sume's. What differs in control, billing and location.
- Compare a local open-weights image model with hosted ones fairly
Same prompt, same shape, no seed: how to test a local Ideogram 4 run against Sume's hosted image models, with a script that prints one image URL per model.
Written by Sume