MAI-Transcribe-2 'one hour in 10 seconds': a short caption clip
Microsoft lists 1hr audio to 10 sec for MAI-Transcribe-2. What that ratio omits for a 20-second clip, and the Sume STT, detach and caption limits that apply.

Microsoft's MAI-Transcribe-2 page lists its end-to-end latency as 1hr audio to 10 sec. That is a batch figure for long audio. It does not tell you what a 20-second clip costs in wall-clock time, and a captioned clip on Sume includes steps no transcription benchmark times: upload, a job queue, and a render. Time your own clip instead.
I read Microsoft's MAI-Transcribe-2 page, the Sume audio detach and video captions docs, and the Sume jobs docs on 2026-10-03. I did not read Sume speed figures anywhere, so I quote none. Treat the Microsoft line as a vendor claim on its own setup, and your own timing as the only number that applies to your pipeline.
What does the 10-second figure cover?
The page puts the 1hr audio to 10 sec latency on MAI-Transcribe-2, and a separate line for the streaming model: about 120 ms to first partial and about 128 ms end to end. The page does not say what hardware, file size or queue conditions sit behind the batch number, and it does not say how it scales to short audio.
| Model on the page | Latency line | Other lines |
|---|---|---|
| MAI-Transcribe-2 | 1hr audio to 10 sec | $0.10 per hour introductory, word-level timestamps, diarization |
| MAI-Transcribe-2-Streaming | ~120 ms first partial, ~128 ms end to end | $0.54 per hour introductory, real-time |
What limits shape a short clip on Sume?
Sume speech-to-text takes a public HTTPS audio URL and accepts up to 10 minutes per request (duration_seconds 1 to 600). Audio detach pulls the track from a hosted video: source up to 1,800 seconds, output up to 900 seconds. A standalone caption job covers video up to 60 seconds. If you do not pass words or script text, captions run speech-to-text for you.
Speech-to-text and audio detach default to asynchronous jobs. You poll a status URL until terminal, then read the result URL. sync mode waits at most 30 seconds, and the docs say that bounds the HTTP wait, not the job.
How do you measure the real number?
Submit your own clip, poll until terminal, and record the wall clock. The script prints seconds, extra polls and word count. Set CLIP_URL and CLIP_SECONDS. For accuracy on the same clip, use the word error rate method.
import json, os, time, urllib.request
def api(method, url, body=None):
req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
return json.load(r)["data"]
start = time.monotonic()
job = api("POST", "https://api.sume.com/v1/stt-1.0/transcribe",
{"audio_url": os.environ["CLIP_URL"], "duration_seconds": int(os.environ["CLIP_SECONDS"]), "mode": "async"})
polls = 0
while not api("GET", job["status_url"])["terminal"]:
polls += 1
time.sleep(1)
result = api("GET", job["result_url"])["result"]
elapsed = time.monotonic() - start
print(f"{elapsed:.1f}s wall clock, {polls} extra polls, {len(result['words'])} words")Sources
Related posts
More in Comparisons
- MAI-Transcribe-2-Streaming partials vs Sume STT word timings
MAI-Transcribe-2-Streaming sends partials after about 100 ms in 60 languages. Sume STT is a job that always returns word timings. Which fits what.
- Make an AI avatar say exact words: Tavus echo mode vs Sume
Tavus echo mode sends text or audio straight to the avatar for playback, skipping perception and speech recognition. Sume's avatar video renders your script.
- Micro-drama lead: Avatar 1.0 or Seedance 2.5? Pick by dialogue
Pick Sume Avatar 1.0 for talking to camera and Seedance 2.5 for a lead who moves through places. Cost per minute: $14.70 against $34.67.
- Clipchamp free captions and silence removal vs a scripted cut list
Clipchamp's pricing page lists AI subtitles and silence removal in the free plan. When a scripted Sume cut list still makes sense, and what each step costs.
Written by Sume