MAI-Transcribe-2 'one hour in 10 seconds': a short caption clip

Microsoft lists 1hr audio to 10 sec for MAI-Transcribe-2. What that ratio omits for a 20-second clip, and the Sume STT, detach and caption limits that apply.

4 min readSume
All posts

Microsoft's MAI-Transcribe-2 page lists its end-to-end latency as 1hr audio to 10 sec. That is a batch figure for long audio. It does not tell you what a 20-second clip costs in wall-clock time, and a captioned clip on Sume includes steps no transcription benchmark times: upload, a job queue, and a render. Time your own clip instead.

I read Microsoft's MAI-Transcribe-2 page, the Sume audio detach and video captions docs, and the Sume jobs docs on 2026-10-03. I did not read Sume speed figures anywhere, so I quote none. Treat the Microsoft line as a vendor claim on its own setup, and your own timing as the only number that applies to your pipeline.

What does the 10-second figure cover?

The page puts the 1hr audio to 10 sec latency on MAI-Transcribe-2, and a separate line for the streaming model: about 120 ms to first partial and about 128 ms end to end. The page does not say what hardware, file size or queue conditions sit behind the batch number, and it does not say how it scales to short audio.

MAI-Transcribe-2 latency lines on Microsoft's page, read 2026-10-03
Model on the pageLatency lineOther lines
MAI-Transcribe-21hr audio to 10 sec$0.10 per hour introductory, word-level timestamps, diarization
MAI-Transcribe-2-Streaming~120 ms first partial, ~128 ms end to end$0.54 per hour introductory, real-time

What limits shape a short clip on Sume?

Sume speech-to-text takes a public HTTPS audio URL and accepts up to 10 minutes per request (duration_seconds 1 to 600). Audio detach pulls the track from a hosted video: source up to 1,800 seconds, output up to 900 seconds. A standalone caption job covers video up to 60 seconds. If you do not pass words or script text, captions run speech-to-text for you.

Speech-to-text and audio detach default to asynchronous jobs. You poll a status URL until terminal, then read the result URL. sync mode waits at most 30 seconds, and the docs say that bounds the HTTP wait, not the job.

How do you measure the real number?

Submit your own clip, poll until terminal, and record the wall clock. The script prints seconds, extra polls and word count. Set CLIP_URL and CLIP_SECONDS. For accuracy on the same clip, use the word error rate method.

import json, os, time, urllib.request

def api(method, url, body=None):
    req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)["data"]

start = time.monotonic()
job = api("POST", "https://api.sume.com/v1/stt-1.0/transcribe",
          {"audio_url": os.environ["CLIP_URL"], "duration_seconds": int(os.environ["CLIP_SECONDS"]), "mode": "async"})
polls = 0
while not api("GET", job["status_url"])["terminal"]:
    polls += 1
    time.sleep(1)
result = api("GET", job["result_url"])["result"]
elapsed = time.monotonic() - start
print(f"{elapsed:.1f}s wall clock, {polls} extra polls, {len(result['words'])} words")

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume