3-hour hearing transcript: MAI Streaming vs Sume STT speaker labels

A 180-minute hearing costs about $1.62 on MAI-Transcribe-2-Streaming and $1.81 on Sume STT with the split, but Sume's result has no speaker labels.

5 min readSume
All posts

A three-hour recorded hearing is cheap to transcribe on either service, but only one of them is documented to label who is speaking. That is 180 audio minutes, about $1.62 on MAI-Transcribe-2-Streaming ($0.54 an audio hour, an introductory rate Microsoft says runs through the end of 2026) and $1.81 on Sume STT 1.0 ($0.01 per audio minute) including the one-cent split job that cuts each recording into pieces of 10 minutes or less.

Inputs are assumptions: one 180-minute English recording with several speakers, transcribed after the fact. Microsoft's rate is from its announcement and Sume's from the API reference, pricing page and timeline audio docs, all read 2026-10-05.

What the bill looks like

Price is not the deciding factor for a hearing: at under two dollars either option is a rounding error against the cost of reviewing the text.

Cost of 180 audio minutes (3.0 hours) across one hearing (read 2026-10-05)
OptionRateTotalJobs
MAI-Transcribe-2-Streaming$0.54 per audio hour (intro rate through end of 2026)$1.62Streaming; no chunking
Sume STT 1.0 (transcription)$0.01 per audio minute ($0.60 an hour)$1.8018 jobs of up to 10 minutes
Sume timeline audio split$0.01 flat per job$0.0101 split jobs, up to 20 ranges each
Sume total$1.81

Speaker attribution is the difference

Microsoft's MAI-Transcribe-2 page lists diarization, timestamps and keyword biasing as built-in features. Sume STT 1.0 returns text, language fields, words[] with start and end times, and optional sentence segments[]. The request fixes provider options such as diarize on the server and rejects them if you send them. Sume STT 1.0 takes duration_seconds from 1 to 600, so a recording longer than 10 minutes is cut into pieces first. The timeline audio split job takes up to 20 ranges, so one split covers up to 200 minutes of audio.

  • With Sume, speaker turns have to come from your own process, for example separate channel recordings per microphone transcribed one by one.
  • A 180-minute recording is 18 ranges of 10 minutes, just under the 20-range limit of one split job.
  • Sentence segments[] give start times you can use as anchors when a reviewer assigns speakers by hand.
  • A machine transcript is a working draft: a legal record needs a human check.

Run it on Sume

This transcribes the whole hearing in 10-minute pieces and prints how many words came back. Each piece is its own job, so one failed piece is retried by its own idempotency key.

The recording URL passed in must already be a media.sume.com audio file, because the split job only reads Sume-hosted audio; import it first.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
def transcribe_long(media_url, total_s, key):
    ranges = [{"start": a, "end": min(a + 600, total_s)} for a in range(0, total_s, 600)]
    split = {"operation": "split", "url": media_url, "ranges": ranges, "output": {"format": "mp3"}}
    parts = run("/v1/timeline-1.0/audio", split, key + "-split")["segments"]
    texts = []
    for i, p in enumerate(parts):
        body = {"audio_url": p["audio_url"], "language_code": "en",
                "duration_seconds": min(600, total_s - 600 * i)}
        texts.append(run("/v1/stt-1.0/transcribe", body, f"{key}-{i}")["text"])
    return " ".join(texts)

text = transcribe_long(os.environ["HEARING_URL"], 180 * 60, "hearing-0412")
print(len(text.split()), "words")

When MAI is the better pick

If the record needs speaker labels out of the box, the MAI-Transcribe-2 feature list is the reason to choose it, and the $0.19 difference is irrelevant. If speakers are recorded on separate channels, or the transcript is only an aid for review, Sume's per-minute job is enough.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume