Weekly 38-minute sermon transcripts: MAI Streaming vs Sume STT cost

A 38-minute sermon every week for a year is 1,976 audio minutes: about $17.78 on MAI-Transcribe-2-Streaming and $20.28 on Sume STT with the split job.

5 min readSume
All posts

A congregation that posts a 38-minute sermon every week for a year needs 52 transcripts. That is 1,976 audio minutes, about $17.78 on MAI-Transcribe-2-Streaming ($0.54 an audio hour, an introductory rate Microsoft says runs through the end of 2026) and $20.28 on Sume STT 1.0 ($0.01 per audio minute) including the one-cent split job that cuts each recording into pieces of 10 minutes or less.

Inputs are assumptions: one 38-minute recording a week for 52 weeks, English, recorded and transcribed after the service. Microsoft's rate is from its announcement and Sume's from the API reference, pricing page and timeline audio docs, all read 2026-10-05.

What the bill looks like

Sume's chunks are the price of a 10-minute job ceiling: each 38-minute sermon becomes four transcription jobs plus one split.

Cost of 1,976 audio minutes (32.9 hours) across 52 weekly recordings (read 2026-10-05)
OptionRateTotalJobs
MAI-Transcribe-2-Streaming$0.54 per audio hour (intro rate through end of 2026)$17.78Streaming; no chunking
Sume STT 1.0 (transcription)$0.01 per audio minute ($0.60 an hour)$19.76208 jobs of up to 10 minutes
Sume timeline audio split$0.01 flat per job$0.52052 split jobs, up to 20 ranges each
Sume total$20.28

A recorded sermon does not need a streaming model

Streaming transcription is built to show words while someone is still talking. If the transcript is published on Monday, the extra machinery has no value, and the choice comes down to price per hour and how much glue code you write. Sume STT 1.0 takes duration_seconds from 1 to 600, so a recording longer than 10 minutes is cut into pieces first. The timeline audio split job takes up to 20 ranges, so one split covers up to 200 minutes of audio.

  • Four chunks of a 38-minute recording are 10, 10, 10 and 8 minutes; pass the real duration_seconds so each job reserves only what it needs.
  • words[] always comes back with start and end in seconds, so each chunk's offsets must be shifted by 600 seconds times its index when you join them.
  • Set language_code when the language is known; omit it for auto-detect.
  • Microsoft's MAI-Transcribe-2 page lists 60+ languages and automatic language detection.

Run it on Sume

This splits a hosted sermon recording, transcribes each piece and joins the text. Use the sermon's length in whole seconds.

The recording URL passed in must already be a media.sume.com audio file, because the split job only reads Sume-hosted audio; import it first.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
def transcribe_long(media_url, total_s, key):
    ranges = [{"start": a, "end": min(a + 600, total_s)} for a in range(0, total_s, 600)]
    split = {"operation": "split", "url": media_url, "ranges": ranges, "output": {"format": "mp3"}}
    parts = run("/v1/timeline-1.0/audio", split, key + "-split")["segments"]
    texts = []
    for i, p in enumerate(parts):
        body = {"audio_url": p["audio_url"], "language_code": "en",
                "duration_seconds": min(600, total_s - 600 * i)}
        texts.append(run("/v1/stt-1.0/transcribe", body, f"{key}-{i}")["text"])
    return " ".join(texts)

print(transcribe_long(os.environ["SERMON_URL"], 38 * 60, "sermon-2026-41")[:300])

When MAI is the better pick

MAI-Transcribe-2-Streaming is $2.50 cheaper over the year in this scenario, and it needs no chunking. If you want live captions during the service, it is the one to use, since Sume STT works on recorded audio. For a recorded archive that also feeds clip cutting and captions, one API key may be worth the price gap.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume