MAI streaming sessions end at one hour: a chunk plan for long audio
A MAI-Transcribe-2-Streaming session lasts at most one hour; Sume STT jobs cap at 10 minutes. Plan overlap-free chunks and shift each chunk's word times.

Neither service takes a three-hour recording in one call. A MAI-Transcribe-2-Streaming session lasts at most one hour, so you reconnect for each hour; Sume STT 1.0 accepts duration_seconds from 1 to 600, so you submit one job per chunk of up to ten minutes and add the chunk's start offset to every word time afterwards.
The two ceilings, side by side
The Microsoft realtime page states the one-hour session limit and that settings cannot change after the first audio append, so each new session restates language and format. On the Sume side the request schema allows duration_seconds 1 to 600 and the audio must be a public HTTPS audio_url.
| Item | MAI-Transcribe-2-Streaming | Sume STT 1.0 |
|---|---|---|
| Unit of work | One WebSocket session | One async job |
| Maximum length | One hour per session | 600 seconds per job |
| Three-hour file | Three sessions | At least 18 jobs |
| Settings between units | Re-sent per session | Re-sent per request |
| Result | delta and committed events | text, words[], optional segments[] |
Cutting without losing words
Cut on silence, not on a fixed clock, or a word straddles two chunks and is transcribed twice or not at all. Record each chunk's start offset in seconds before you upload it. Sume returns words[] as {word, start, end} with times relative to the chunk, so the merge is arithmetic.
import json, sys
def merge(chunks):
"""chunks: list of (offset_seconds, stt_result_dict)"""
out = []
for offset, result in chunks:
for w in result.get("words", []):
out.append({"word": w["word"],
"start": round(w["start"] + offset, 3),
"end": round(w["end"] + offset, 3)})
return out
if __name__ == "__main__":
demo = [
(0, {"words": [{"word": "Hello.", "start": 0.1, "end": 0.5}]}),
(600, {"words": [{"word": "Again.", "start": 0.2, "end": 0.7}]}),
]
json.dump(merge(demo), sys.stdout, indent=1)
A reconnect and chunk checklist
Whichever service you use, long audio turns into a bookkeeping problem. These habits keep it manageable.
- Store the offset of every chunk next to its job id before you submit, not after the results come back.
- Cut at the quietest point within a few seconds of the target length, so no word is split.
- Submit chunks in parallel with a small cap, then sort by offset when merging.
- Keep each chunk under 600 seconds and pass the real length as
duration_seconds, which lets the reservation match the audio. - Re-run only the failed chunk rather than the whole file.
Limits and a caution
Each Sume result caps words[] at 20000 entries and reports words_truncated if it ever hits that, which a 600 second chunk stays well under. If you merge chunks yourself, keep the final merged list outside any single job's cap.
Not every case needs chunking. If you have a 12 minute file, two jobs of 6 minutes is simpler than a socket, and the per-job duration_seconds hint also makes the usage reservation match what you need. For live audio longer than an hour, plan the reconnect: the preview has no SLA, so the reconnect path is the part to test first.
Sources
Related posts
More in Developers
- MAI Flash 429 and spend limits vs Sume queued jobs and idempotency
OpenRouter's MAI-Voice-2.1-Flash returns 429 on rate limits and has spend limits. Sume accepts jobs into queued status. See an idempotent double submit.
- MAI Flash defaults to PCM: wrap it as WAV, vs Sume output formats
OpenRouter lists MAI-Voice-2.1-Flash output as mp3 or pcm, default pcm, which will not play as saved. Wrap it as WAV in Python; Sume defaults to mp3.
- MAI-Voice 24 kHz 160 kbps mp3 header vs Sume output_format settings
Microsoft's REST example asks for audio-24khz-160kbitrate-mono-mp3. Sume has no 160 kbps option: mp3 bit rates are 32, 64, 96, 128 or 192 kbps at up to 48 kHz.
- MAI voice ids end in a model name; Sume voice ids do not
A MAI id like en-US-Harper:MAI-Voice-2.1-Flash is not a Sume voice id. Sume takes a UUID or voi_ plus 32 hex and returns 400 invalid_voice_id for anything else.
Written by Sume