A 'Pareto frontier' speech-to-text claim: plot your own cost and delay
Microsoft says MAI-Transcribe-2-Streaming sits on the Pareto frontier. A frontier needs two axes and your data. Plot Sume STT cost against measured job time.

A Pareto frontier claim means no other model is better on one axis without being worse on another. Microsoft's October 1 announcement uses that language for MAI-Transcribe-2-Streaming, alongside a first-ranked claim on Artificial Analysis, first hypotheses in just over 100 ms, 60 languages and an introductory price of $0.54 per hour through the end of the year. To test a frontier claim for your use, you need your own two axes. For a file-based service like Sume STT they are cost per audio minute and job time per clip.
What a frontier needs
Take every option you would actually consider, put cost on one axis and delay on the other, and mark an option as dominated if another is both cheaper and faster. What is left is your frontier. The vendor's chart answers a different question, because it uses their test set, their delay definition and their comparison group. For streaming models, delay means time to first hypothesis. For a file job it means submit to result, which is a different number and not comparable.
| Option | Cost axis | Delay axis |
|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per hour (intro, through year end) = $0.009 per minute | First hypotheses just over 100 ms (vendor figure) |
| Sume STT 1.0 | $0.01 per audio minute = $0.60 per hour | Job time, measured by you |
| Your own candidates | Price page, dated | Measured on your clips |
Measure Sume's side
Time ten clips of one minute each, using duration_seconds: 60 so the reservation matches. Record the seconds from submit to the terminal status. Cost is fixed at the published rate, so ten clips cost 10 x $0.01 = 10 cents. The median and the slowest run give you the delay axis.
import os, statistics, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
url = "https://example.com/clip-60s.mp3"
secs = []
for _ in range(10):
t0 = time.time()
d = requests.post("https://api.sume.com/v1/stt-1.0/transcribe", headers=H,
json={"audio_url": url, "duration_seconds": 60}, timeout=60).json()["data"]
while not d.get("terminal"):
time.sleep(1)
d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
secs.append(time.time() - t0)
print("median", round(statistics.median(secs), 1), "max", round(max(secs), 1))Read the result honestly
Sume STT will not appear on a streaming-latency axis, because it is not a stream. It belongs on a file-transcription frontier, where the question is cost and turnaround for finished audio. If your product needs words while someone is still speaking, the claim is relevant and Sume is not the answer; if your product transcribes recordings, the introductory $0.009 per minute is a price difference of 0.1 cent per minute, or 6 cents per hour, against $0.01.
Remember the intro price ends after this year, and the page does not say what replaces it. Do not build a multi-year model on it.
Before you draw the chart
Decide which delay you care about. A meeting summary tool can wait minutes; a captioning overlay on a live stream cannot. If you wait minutes anyway, the delay axis collapses and cost, language coverage and accuracy on your audio decide the choice. Add those as extra columns, so a dominated option on cost and delay can still win on a dimension the frontier ignores, such as a language you need.
Record the date you read every price. The Microsoft figure is an introductory rate that the announcement ties to the end of the year, so a chart built today expires in January.
Sources
Related posts
More in Comparisons
- File-size limits as average bitrate for a 60-second clip
A 60-second clip can average 3.3 Mbps on free ArtStation, 6.7 on a Linktree background, 25 on Reels and 60 on Truth Social. Full table and script.
- Longest clip per platform and how many Sume trim jobs it takes
Truth Social 15 minutes, Reels API 15, TikTok API 10, Kick and Shorts 3, ArtStation 1. How many trim jobs a 30-minute source needs for each, and the cost.
- Predis Core: 325 images or 33 videos vs per-image pricing on Sume
Predis Core lists 1,300 credits, about 325 images or 33 videos. Divide the plan price and compare to Sume's per-image estimate of $0.2225.
- Price per finished video: Predis, Submagic, Zebracat vs Sume
Dividing plan prices by listed video counts gives $0.69 to $2.60 per video. A Sume captioned 30-second render is $0.32; an avatar ad is $5.52.
Written by Sume