Live captions for a 90-minute match: streaming STT, then Sume replay
A 90-minute match costs about $0.81 on MAI-Transcribe-2-Streaming. Sume STT works on recordings: $0.91 after the match for replay captions.

Use a streaming model such as MAI-Transcribe-2-Streaming for captions during the match, and Sume STT 1.0 for the replay video afterwards. A 90-minute match is 1.5 audio hours: about $0.81 at Microsoft's $0.54 an hour, and $0.91 on Sume including the one-cent split job, because Sume transcribes recorded audio in jobs of up to 10 minutes rather than live.
Microsoft's announcement says the streaming model delivers first hypotheses in just over 100 ms and revises partial text as context arrives. Prices are from that announcement, the Sume API reference and the timeline audio docs, read 2026-10-05.
Two jobs, two tools
Live captions and replay captions solve different problems. Live captions must appear within a fraction of a second and can be corrected later. Replay captions can wait for the match to finish and should be final, timed and ready to burn into the video.
| Need | Tool | Cost for 90 minutes | What you get |
|---|---|---|---|
| Captions during the match | MAI-Transcribe-2-Streaming ($0.54 per audio hour) | $0.81 | Partial hypotheses that revise, in 60 languages per Microsoft |
| Captions on the replay | Sume STT 1.0 ($0.01 per audio minute) plus a $0.01 split job | $0.91 | Final text, words[] timings and sentence segments[] per 10-minute chunk |
Why the replay should be re-transcribed
Partial hypotheses change as context arrives, so a live caption log is not the final text of the match. A second pass over the recording gives one stable transcript with word timings.
- Sume STT takes
duration_secondsfrom 1 to 600, so the recording is split into up to nine 10-minute pieces. - Each piece's timings start at zero: add 600 seconds times the piece index to line them up with the full match.
- Request
segmentation: {"mode": "sentence"}to get one start and end time per caption cue. - Burn only the final cues into the replay clip.
Run it on Sume
This splits the recorded match into 10-minute pieces, transcribes each with sentence segmentation and prints cue start times on the full-match timeline. The source must be a media.sume.com audio file.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
total = 90 * 60
ranges = [{"start": a, "end": min(a + 600, total)} for a in range(0, total, 600)]
split = {"operation": "split", "url": os.environ["MATCH_URL"], "ranges": ranges}
parts = run("/v1/timeline-1.0/audio", split, "match-split")["segments"]
for i, p in enumerate(parts):
body = {"audio_url": p["audio_url"], "duration_seconds": 600,
"segmentation": {"mode": "sentence"}}
for seg in run("/v1/stt-1.0/transcribe", body, f"match-{i}")["segments"]:
print(round(seg["start"] + 600 * i, 1), seg["text"])
When Sume is not the tool
If you need words on screen while the match is being played, Sume STT does not do that: it returns a result when the job finishes. Use a streaming model for the live layer and use Sume for the clean replay.
Sources
Related posts
More in Use cases
- Live-selling replay: 12 clips in 3 languages, captions for $7.68
Trim 12 product clips from a live-selling replay and caption them in three languages on Sume for $7.68: $0.04 a clip to prepare, $0.20 per caption job.
- Localize a 30-second holiday ad into 8 languages for about $2.65
Eight voice-over versions of a 30-second holiday ad cost $2.65 in Sume usage: TTS $0.03, Timeline $0.10, captions $0.20 per language, plus $0.01 to transcribe.
- Localize a 30-second product video into 3 languages: swap the spine
Keep one edit and swap only the voice: three Timeline renders ($0.10) plus three caption jobs ($0.20) cost $0.90 on Sume for English, Spanish and French.
- Localize app store screenshot captions: 6 locales x 5 shots, $1.13
Translate the caption on 5 marketing screenshots into 6 locales with Ideogram 4.5 edits on Sume: 30 calls, $1.125 at low, and why real UI gets re-captured.
Written by Sume