Do you need streaming transcription? A recorded clip needs a batch job

MAI-Transcribe-2-Streaming is built for live audio. For a finished clip, Sume STT at $0.01 a minute returns text and word timings. How to decide which you need.

4 min readSume
All posts

You need streaming transcription only when text must appear while someone is still talking. For a recorded clip, a batch job is the simpler fit: Sume STT returns the whole transcript with word timings for $0.01 per audio minute, and a file up to 10 minutes goes in one job (API reference). Microsoft's MAI-Transcribe-2-Streaming, launched October 1, 2026, is the other kind: first partial results arrive in just over 100 ms, in 60 languages, at $0.54 per hour of audio through the end of the year (Microsoft AI, read 2026-10-04).

Which one fits which job

Streaming versus batch speech to text, from the vendor page and the Sume contract (read 2026-10-04)
JobStreaming modelSume batch STT
Live captions on a callFitsDoes not fit
Voice agent hearing a callerFitsDoes not fit
Subtitles for a finished shortWorks, but partials are not neededFits
Script from a recorded videoWorks, but partials are not neededFits
Word timings for captionsNot described on the pageAlways returned

The price per minute

Microsoft's $0.54 per hour is $0.009 per minute, and Sume's $0.01 per minute is $0.60 per hour. The Microsoft rate is introductory and runs through year end, so it can change. For 100 recorded clips of 90 seconds each, 150 minutes of audio, Sume's rate comes to $1.50.

What the batch job does not do

Sume STT rejects diarize and tag_audio_events, so there are no speaker labels or event tags. A single job takes at most 10 minutes of audio, and duration_seconds should be sent as a hint from 1 to 600 so the reservation matches the clip. A sync call waits at most 30 seconds, so submit async or webhook for longer clips and poll as the jobs docs describe.

Decision rule

Ask one question: does anyone act on the text before the audio ends? If yes, use a streaming model. If the text is read after the recording ends, which is the case for captions, show notes and search, batch is enough and the contract is simpler. See batch transcription for the multi-file pattern.

Sources

Related posts

More in Models

All Models posts

Written by Sume