MAI-Transcribe-2-Streaming or a batch STT job: which fits?

MAI-Transcribe-2-Streaming returns first partials in just over 100 ms. Sume's speech-to-text is a batch job up to 10 minutes at $0.01 per minute.

5 min readSume
All posts

Use a streaming model such as MAI-Transcribe-2-Streaming when text must appear while someone is still speaking. Use a batch job like Sume's speech-to-text when you have a finished recording and want a transcript with sentence timings. Microsoft reports first partials in just over 100 ms and 60 languages for MAI-Transcribe-2-Streaming.

Streaming versus batch

Streaming and batch transcription (read 2026-10-04)
ItemMAI-Transcribe-2-StreamingSume STT
ModeStreamingJob on a public audio_url
Price$0.54 per hour through year end$0.01 per audio minute ($0.60 per hour)
Latency claimFirst partials just over 100 msNot claimed
Max inputContinuous stream10 minutes (600 s) per job
OutputPartial and final textSentence segmentation

Reported accuracy

Unite.ai reports 2.5% final word error rate, 2.8% for the first partial and 0.13 seconds to the final transcript. Those are vendor-reported; I ran no test.

Batch with Sume

Send audio_url, an optional language_code and duration_seconds from 1 to 600. Sentence segmentation takes boundary_lead_ms from 0 to 500 with a default of 70.

  • diarize and tag_audio_events are fixed on the server and a request that sends them returns 400.
  • Longer audio: split with Timeline audio first, $0.01 per job.
  • Check the video has an audio track before transcribing.

Where each is available

MAI-Transcribe-2-Streaming is listed on Foundry, MAI Playground, Vercel, OpenRouter and Azure Voice Live, with LiveKit coming soon.

Sources

Related posts

More in Models

All Models posts

Written by Sume