Twelve 45-minute user interviews transcribed: MAI Streaming vs Sume

Twelve 45-minute research interviews are 540 audio minutes: about $4.86 on MAI-Transcribe-2-Streaming and $5.52 on Sume STT, with five jobs per interview.

5 min readSume
All posts

A research round of twelve 45-minute user interviews is a small transcription job. That is 540 audio minutes, about $4.86 on MAI-Transcribe-2-Streaming ($0.54 an audio hour, an introductory rate Microsoft says runs through the end of 2026) and $5.52 on Sume STT 1.0 ($0.01 per audio minute) including the one-cent split job that cuts each recording into pieces of 10 minutes or less.

Inputs are assumptions: twelve English interviews of 45 minutes, transcribed once after the sessions. Microsoft's rate is from its announcement and Sume's from the API reference, pricing page and timeline audio docs, all read 2026-10-05.

What the bill looks like

Sume turns each interview into five transcription jobs (four of 10 minutes and one of 5), so a round is 60 jobs.

Cost of 540 audio minutes (9.0 hours) across 12 interviews (read 2026-10-05)
OptionRateTotalJobs
MAI-Transcribe-2-Streaming$0.54 per audio hour (intro rate through end of 2026)$4.86Streaming; no chunking
Sume STT 1.0 (transcription)$0.01 per audio minute ($0.60 an hour)$5.4060 jobs of up to 10 minutes
Sume timeline audio split$0.01 flat per job$0.12012 split jobs, up to 20 ranges each
Sume total$5.52

What a research team does with the output

Interview transcripts are mostly searched and quoted, so timestamps matter more than live speed. Sume STT 1.0 takes duration_seconds from 1 to 600, so a recording longer than 10 minutes is cut into pieces first. The timeline audio split job takes up to 20 ranges, so one split covers up to 200 minutes of audio.

  • words[] timings let a reviewer jump to a quote in the recording.
  • Add 600 seconds times the chunk index to each chunk's timings before you publish a single timeline.
  • Use segmentation: {"mode": "sentence"} on the transcription request to get sentence segments with start and end times for quote cards.
  • Interviews with two speakers need a manual or per-channel step for attribution; the Sume request does not accept a diarize flag.

Run it on Sume

This transcribes each interview file in the folder list and writes a text file per interview. The input is a list of hosted audio URLs and each is 45 minutes.

The recording URL passed in must already be a media.sume.com audio file, because the split job only reads Sume-hosted audio; import it first.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
def transcribe_long(media_url, total_s, key):
    ranges = [{"start": a, "end": min(a + 600, total_s)} for a in range(0, total_s, 600)]
    split = {"operation": "split", "url": media_url, "ranges": ranges, "output": {"format": "mp3"}}
    parts = run("/v1/timeline-1.0/audio", split, key + "-split")["segments"]
    texts = []
    for i, p in enumerate(parts):
        body = {"audio_url": p["audio_url"], "language_code": "en",
                "duration_seconds": min(600, total_s - 600 * i)}
        texts.append(run("/v1/stt-1.0/transcribe", body, f"{key}-{i}")["text"])
    return " ".join(texts)

urls = os.environ["INTERVIEW_URLS"].split(",")
for n, url in enumerate(urls, 1):
    text = transcribe_long(url, 45 * 60, f"interview-{n:02d}")
    open(f"interview-{n:02d}.txt", "w", encoding="utf-8").write(text)

When MAI is the better pick

MAI-Transcribe-2-Streaming is $0.66 cheaper and takes no chunking. It also lists diarization on the MAI-Transcribe-2 page, which a two-speaker interview benefits from. Choose Sume if the quotes then go straight into cut-down clips and captions in the same account.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume