Forty 55-minute earnings calls transcribed: MAI Streaming vs Sume STT
Forty earnings calls of 55 minutes are 2,200 audio minutes: about $19.80 on MAI-Transcribe-2-Streaming and $22.40 on Sume STT including the split jobs.

Transcribing the earnings calls of forty companies in one quarter is 2,200 minutes of audio. That is 2,200 audio minutes, about $19.80 on MAI-Transcribe-2-Streaming ($0.54 an audio hour, an introductory rate Microsoft says runs through the end of 2026) and $22.40 on Sume STT 1.0 ($0.01 per audio minute) including the one-cent split job that cuts each recording into pieces of 10 minutes or less.
Inputs are assumptions: forty 55-minute English calls in a quarter. Microsoft's rate is from its announcement and Sume's from the API reference, pricing page and timeline audio docs, all read 2026-10-05.
What the bill looks like
Each call is six transcription jobs on Sume (five of 10 minutes and one of 5), so a quarter is 240 jobs plus 40 splits.
| Option | Rate | Total | Jobs |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per audio hour (intro rate through end of 2026) | $19.80 | Streaming; no chunking |
| Sume STT 1.0 (transcription) | $0.01 per audio minute ($0.60 an hour) | $22.00 | 240 jobs of up to 10 minutes |
| Sume timeline audio split | $0.01 flat per job | $0.400 | 40 split jobs, up to 20 ranges each |
| Sume total | $22.40 |
Names and numbers decide whether the transcript is usable
Calls are full of tickers, product names and figures. Microsoft's MAI-Transcribe-2 page lists keyword biasing for domain terminology. Sume STT 1.0 has no keyword list on the request: it takes audio_url, language_code, duration_seconds, optional segmentation and metadata. Sume STT 1.0 takes duration_seconds from 1 to 600, so a recording longer than 10 minutes is cut into pieces first. The timeline audio split job takes up to 20 ranges, so one split covers up to 200 minutes of audio.
- Correct company names in a second pass, using the
words[]timings to find them. - Pass
metadatasuch as the ticker and quarter; it is stored with the job request and not sent to the provider. - Use a stable idempotency key per call and chunk, such as
call-q3-ACME-2, so a rerun is safe. - The recordings must be reachable at a public HTTPS URL; Sume prefers its own media host.
Run it on Sume
This transcribes each call from a list of hosted audio URLs and writes one text file per call.
The recording URL passed in must already be a media.sume.com audio file, because the split job only reads Sume-hosted audio; import it first.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
def transcribe_long(media_url, total_s, key):
ranges = [{"start": a, "end": min(a + 600, total_s)} for a in range(0, total_s, 600)]
split = {"operation": "split", "url": media_url, "ranges": ranges, "output": {"format": "mp3"}}
parts = run("/v1/timeline-1.0/audio", split, key + "-split")["segments"]
texts = []
for i, p in enumerate(parts):
body = {"audio_url": p["audio_url"], "language_code": "en",
"duration_seconds": min(600, total_s - 600 * i)}
texts.append(run("/v1/stt-1.0/transcribe", body, f"{key}-{i}")["text"])
return " ".join(texts)
calls = os.environ["CALL_URLS"].split(",")
for i, url in enumerate(calls, 1):
text = transcribe_long(url, 55 * 60, f"call-q3-{i:03d}")
open(f"call-{i:03d}.txt", "w", encoding="utf-8").write(text)
When MAI is the better pick
For finance vocabulary, keyword biasing is the feature to weigh: it is listed on the MAI-Transcribe-2 page and it is not in the Sume request. The price gap of $2.60 per quarter is small next to that difference.
Sources
Related posts
More in Pricing
- Free AI video generator: what the Sume Free plan includes
The Sume Free plan is for trying image generation. Video is paid per clip from the wallet; the cheapest clips cost about 12 to 13 cents.
- Free plan vs paid wallet for AI video: what each one buys on Sume
On Sume the plan sets how many video jobs run at once and your request rates; the wallet pays for each clip. A top-up never raises concurrency.
- Cheapest TTS API this week: Gemini Flash-Lite, $0.0015 per 10 seconds
Gemini 3.8 Flash-Lite TTS is about $0.0015 per 10 seconds of audio, about $0.54 an hour. Flash is $0.00225. Sume's Sonic route is about 4x that per hour.
- Gemini image prices per million tokens: Pro $120, NB2 $60, Lite $30
Gemini image output costs $120 per 1M tokens on Pro, $60 on Nano Banana 2 and $30 on Lite. Divide by 1M to price one image; Sume bills a flat per-image rate.
Written by Sume