A 'Pareto frontier' speech-to-text claim: plot your own cost and delay

Microsoft says MAI-Transcribe-2-Streaming sits on the Pareto frontier. A frontier needs two axes and your data. Plot Sume STT cost against measured job time.

5 min readSume
All posts

A Pareto frontier claim means no other model is better on one axis without being worse on another. Microsoft's October 1 announcement uses that language for MAI-Transcribe-2-Streaming, alongside a first-ranked claim on Artificial Analysis, first hypotheses in just over 100 ms, 60 languages and an introductory price of $0.54 per hour through the end of the year. To test a frontier claim for your use, you need your own two axes. For a file-based service like Sume STT they are cost per audio minute and job time per clip.

What a frontier needs

Take every option you would actually consider, put cost on one axis and delay on the other, and mark an option as dominated if another is both cheaper and faster. What is left is your frontier. The vendor's chart answers a different question, because it uses their test set, their delay definition and their comparison group. For streaming models, delay means time to first hypothesis. For a file job it means submit to result, which is a different number and not comparable.

Axes for a speech-to-text frontier, Microsoft announcement and Sume STT docs, read 2026-10-05
OptionCost axisDelay axis
MAI-Transcribe-2-Streaming$0.54 per hour (intro, through year end) = $0.009 per minuteFirst hypotheses just over 100 ms (vendor figure)
Sume STT 1.0$0.01 per audio minute = $0.60 per hourJob time, measured by you
Your own candidatesPrice page, datedMeasured on your clips

Measure Sume's side

Time ten clips of one minute each, using duration_seconds: 60 so the reservation matches. Record the seconds from submit to the terminal status. Cost is fixed at the published rate, so ten clips cost 10 x $0.01 = 10 cents. The median and the slowest run give you the delay axis.

import os, statistics, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
url = "https://example.com/clip-60s.mp3"
secs = []
for _ in range(10):
    t0 = time.time()
    d = requests.post("https://api.sume.com/v1/stt-1.0/transcribe", headers=H,
        json={"audio_url": url, "duration_seconds": 60}, timeout=60).json()["data"]
    while not d.get("terminal"):
        time.sleep(1)
        d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
    secs.append(time.time() - t0)
print("median", round(statistics.median(secs), 1), "max", round(max(secs), 1))

Read the result honestly

Sume STT will not appear on a streaming-latency axis, because it is not a stream. It belongs on a file-transcription frontier, where the question is cost and turnaround for finished audio. If your product needs words while someone is still speaking, the claim is relevant and Sume is not the answer; if your product transcribes recordings, the introductory $0.009 per minute is a price difference of 0.1 cent per minute, or 6 cents per hour, against $0.01.

Remember the intro price ends after this year, and the page does not say what replaces it. Do not build a multi-year model on it.

Before you draw the chart

Decide which delay you care about. A meeting summary tool can wait minutes; a captioning overlay on a live stream cannot. If you wait minutes anyway, the delay axis collapses and cost, language coverage and accuracy on your audio decide the choice. Add those as extra columns, so a dominated option on cost and delay can still win on a dimension the frontier ignores, such as a language you need.

Record the date you read every price. The Microsoft figure is an introductory rate that the announcement ties to the end of the year, so a chart built today expires in January.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume