Do you need streaming transcription? A recorded clip needs a batch job
MAI-Transcribe-2-Streaming is built for live audio. For a finished clip, Sume STT at $0.01 a minute returns text and word timings. How to decide which you need.

You need streaming transcription only when text must appear while someone is still talking. For a recorded clip, a batch job is the simpler fit: Sume STT returns the whole transcript with word timings for $0.01 per audio minute, and a file up to 10 minutes goes in one job (API reference). Microsoft's MAI-Transcribe-2-Streaming, launched October 1, 2026, is the other kind: first partial results arrive in just over 100 ms, in 60 languages, at $0.54 per hour of audio through the end of the year (Microsoft AI, read 2026-10-04).
Which one fits which job
| Job | Streaming model | Sume batch STT |
|---|---|---|
| Live captions on a call | Fits | Does not fit |
| Voice agent hearing a caller | Fits | Does not fit |
| Subtitles for a finished short | Works, but partials are not needed | Fits |
| Script from a recorded video | Works, but partials are not needed | Fits |
| Word timings for captions | Not described on the page | Always returned |
The price per minute
Microsoft's $0.54 per hour is $0.009 per minute, and Sume's $0.01 per minute is $0.60 per hour. The Microsoft rate is introductory and runs through year end, so it can change. For 100 recorded clips of 90 seconds each, 150 minutes of audio, Sume's rate comes to $1.50.
What the batch job does not do
Sume STT rejects diarize and tag_audio_events, so there are no speaker labels or event tags. A single job takes at most 10 minutes of audio, and duration_seconds should be sent as a hint from 1 to 600 so the reservation matches the clip. A sync call waits at most 30 seconds, so submit async or webhook for longer clips and poll as the jobs docs describe.
Decision rule
Ask one question: does anyone act on the text before the audio ends? If yes, use a streaming model. If the text is read after the recording ends, which is the case for captions, show notes and search, batch is enough and the contract is simpler. See batch transcription for the multi-file pattern.
Sources
Related posts
More in Models
- Does Lyria 3.5 audio carry a SynthID watermark?
Yes: Google's Gemini API music docs say Lyria audio carries a SynthID watermark. The Sume docs I read do not describe watermarking either way.
- Draw boxes on a photo, then edit it: Sketch-style markup on Sume
ChatGPT Sketch is a drawing panel inside ChatGPT. On the API you get the same effect by sending a marked-up photo as a reference and saying what the marks mean.
- Eleven v4 stacked tags: direct emotion without them
ElevenLabs v4 adds stackable expression tags and 10-second cloning. Sume's docs list no clone route; here is how to direct a voice via tts_create.
- eleven_v4 vs eleven_v4_turbo: model IDs, endpoints, which to pick
ElevenLabs lists eleven_v4 for expressive speech with cloning in 90+ languages and eleven_v4_turbo at about 100 ms median latency. Which fits a video pipeline.
Written by Sume