Live captions for a 90-minute match: streaming STT, then Sume replay

A 90-minute match costs about $0.81 on MAI-Transcribe-2-Streaming. Sume STT works on recordings: $0.91 after the match for replay captions.

5 min readSume
All posts

Use a streaming model such as MAI-Transcribe-2-Streaming for captions during the match, and Sume STT 1.0 for the replay video afterwards. A 90-minute match is 1.5 audio hours: about $0.81 at Microsoft's $0.54 an hour, and $0.91 on Sume including the one-cent split job, because Sume transcribes recorded audio in jobs of up to 10 minutes rather than live.

Microsoft's announcement says the streaming model delivers first hypotheses in just over 100 ms and revises partial text as context arrives. Prices are from that announcement, the Sume API reference and the timeline audio docs, read 2026-10-05.

Two jobs, two tools

Live captions and replay captions solve different problems. Live captions must appear within a fraction of a second and can be corrected later. Replay captions can wait for the match to finish and should be final, timed and ready to burn into the video.

Live captions vs replay captions for a 90-minute match (read 2026-10-05)
NeedToolCost for 90 minutesWhat you get
Captions during the matchMAI-Transcribe-2-Streaming ($0.54 per audio hour)$0.81Partial hypotheses that revise, in 60 languages per Microsoft
Captions on the replaySume STT 1.0 ($0.01 per audio minute) plus a $0.01 split job$0.91Final text, words[] timings and sentence segments[] per 10-minute chunk

Why the replay should be re-transcribed

Partial hypotheses change as context arrives, so a live caption log is not the final text of the match. A second pass over the recording gives one stable transcript with word timings.

  • Sume STT takes duration_seconds from 1 to 600, so the recording is split into up to nine 10-minute pieces.
  • Each piece's timings start at zero: add 600 seconds times the piece index to line them up with the full match.
  • Request segmentation: {"mode": "sentence"} to get one start and end time per caption cue.
  • Burn only the final cues into the replay clip.

Run it on Sume

This splits the recorded match into 10-minute pieces, transcribes each with sentence segmentation and prints cue start times on the full-match timeline. The source must be a media.sume.com audio file.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]

total = 90 * 60
ranges = [{"start": a, "end": min(a + 600, total)} for a in range(0, total, 600)]
split = {"operation": "split", "url": os.environ["MATCH_URL"], "ranges": ranges}
parts = run("/v1/timeline-1.0/audio", split, "match-split")["segments"]
for i, p in enumerate(parts):
    body = {"audio_url": p["audio_url"], "duration_seconds": 600,
            "segmentation": {"mode": "sentence"}}
    for seg in run("/v1/stt-1.0/transcribe", body, f"match-{i}")["segments"]:
        print(round(seg["start"] + 600 * i, 1), seg["text"])

When Sume is not the tool

If you need words on screen while the match is being played, Sume STT does not do that: it returns a result when the job finishes. Use a streaming model for the live layer and use Sume for the clean replay.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume