3-hour hearing transcript: MAI Streaming vs Sume STT speaker labels
A 180-minute hearing costs about $1.62 on MAI-Transcribe-2-Streaming and $1.81 on Sume STT with the split, but Sume's result has no speaker labels.

A three-hour recorded hearing is cheap to transcribe on either service, but only one of them is documented to label who is speaking. That is 180 audio minutes, about $1.62 on MAI-Transcribe-2-Streaming ($0.54 an audio hour, an introductory rate Microsoft says runs through the end of 2026) and $1.81 on Sume STT 1.0 ($0.01 per audio minute) including the one-cent split job that cuts each recording into pieces of 10 minutes or less.
Inputs are assumptions: one 180-minute English recording with several speakers, transcribed after the fact. Microsoft's rate is from its announcement and Sume's from the API reference, pricing page and timeline audio docs, all read 2026-10-05.
What the bill looks like
Price is not the deciding factor for a hearing: at under two dollars either option is a rounding error against the cost of reviewing the text.
| Option | Rate | Total | Jobs |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per audio hour (intro rate through end of 2026) | $1.62 | Streaming; no chunking |
| Sume STT 1.0 (transcription) | $0.01 per audio minute ($0.60 an hour) | $1.80 | 18 jobs of up to 10 minutes |
| Sume timeline audio split | $0.01 flat per job | $0.010 | 1 split jobs, up to 20 ranges each |
| Sume total | $1.81 |
Speaker attribution is the difference
Microsoft's MAI-Transcribe-2 page lists diarization, timestamps and keyword biasing as built-in features. Sume STT 1.0 returns text, language fields, words[] with start and end times, and optional sentence segments[]. The request fixes provider options such as diarize on the server and rejects them if you send them. Sume STT 1.0 takes duration_seconds from 1 to 600, so a recording longer than 10 minutes is cut into pieces first. The timeline audio split job takes up to 20 ranges, so one split covers up to 200 minutes of audio.
- With Sume, speaker turns have to come from your own process, for example separate channel recordings per microphone transcribed one by one.
- A 180-minute recording is 18 ranges of 10 minutes, just under the 20-range limit of one split job.
- Sentence
segments[]give start times you can use as anchors when a reviewer assigns speakers by hand. - A machine transcript is a working draft: a legal record needs a human check.
Run it on Sume
This transcribes the whole hearing in 10-minute pieces and prints how many words came back. Each piece is its own job, so one failed piece is retried by its own idempotency key.
The recording URL passed in must already be a media.sume.com audio file, because the split job only reads Sume-hosted audio; import it first.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
def transcribe_long(media_url, total_s, key):
ranges = [{"start": a, "end": min(a + 600, total_s)} for a in range(0, total_s, 600)]
split = {"operation": "split", "url": media_url, "ranges": ranges, "output": {"format": "mp3"}}
parts = run("/v1/timeline-1.0/audio", split, key + "-split")["segments"]
texts = []
for i, p in enumerate(parts):
body = {"audio_url": p["audio_url"], "language_code": "en",
"duration_seconds": min(600, total_s - 600 * i)}
texts.append(run("/v1/stt-1.0/transcribe", body, f"{key}-{i}")["text"])
return " ".join(texts)
text = transcribe_long(os.environ["HEARING_URL"], 180 * 60, "hearing-0412")
print(len(text.split()), "words")
When MAI is the better pick
If the record needs speaker labels out of the box, the MAI-Transcribe-2 feature list is the reason to choose it, and the $0.19 difference is irrelevant. If speakers are recorded on separate channels, or the transcript is only an aid for review, Sume's per-minute job is enough.
Sources
Related posts
More in Comparisons
- Wan 3.0, H3 and H3 Max tie at $0.0625 a second at 480p on Sume
At 480p, wan-3.0, minimax-h3 and minimax-h3-max all bill $0.0625 a second on Sume, $0.63 for 10 seconds. Break the tie on length and tiers above.
- Does TikTok auto-label AI video? Only its own AI effects
TikTok labels content made with its official AI effects automatically. A Sume clip is made outside the app, so you disclose it yourself when it looks real.
- Transcribe 500 hours: Speechmatics tiers vs Sume STT cost
500 hours of audio: $60 on Speechmatics Melia 1, $115 Standard, $190 Enhanced, $300 on Sume STT 1.0. The math, the 10-minute job limit, free credit.
- Translate packaging text: Ideogram 4.5 vs Nano Banana Pro per language
Translate label text in one image per language. On Sume Ideogram 4.5 costs $0.0375 to $0.275 per edit; Nano Banana Pro $0.1875 at 1K to 2K. Limits compared.
Written by Sume