MAI streaming sessions end at one hour: a chunk plan for long audio

A MAI-Transcribe-2-Streaming session lasts at most one hour; Sume STT jobs cap at 10 minutes. Plan overlap-free chunks and shift each chunk's word times.

4 min readSume
All posts

Neither service takes a three-hour recording in one call. A MAI-Transcribe-2-Streaming session lasts at most one hour, so you reconnect for each hour; Sume STT 1.0 accepts duration_seconds from 1 to 600, so you submit one job per chunk of up to ten minutes and add the chunk's start offset to every word time afterwards.

The two ceilings, side by side

The Microsoft realtime page states the one-hour session limit and that settings cannot change after the first audio append, so each new session restates language and format. On the Sume side the request schema allows duration_seconds 1 to 600 and the audio must be a public HTTPS audio_url.

Long-audio limits (Microsoft realtime page read 2026-10-05; Sume from the repo)
ItemMAI-Transcribe-2-StreamingSume STT 1.0
Unit of workOne WebSocket sessionOne async job
Maximum lengthOne hour per session600 seconds per job
Three-hour fileThree sessionsAt least 18 jobs
Settings between unitsRe-sent per sessionRe-sent per request
Resultdelta and committed eventstext, words[], optional segments[]

Cutting without losing words

Cut on silence, not on a fixed clock, or a word straddles two chunks and is transcribed twice or not at all. Record each chunk's start offset in seconds before you upload it. Sume returns words[] as {word, start, end} with times relative to the chunk, so the merge is arithmetic.

import json, sys

def merge(chunks):
    """chunks: list of (offset_seconds, stt_result_dict)"""
    out = []
    for offset, result in chunks:
        for w in result.get("words", []):
            out.append({"word": w["word"],
                        "start": round(w["start"] + offset, 3),
                        "end": round(w["end"] + offset, 3)})
    return out

if __name__ == "__main__":
    demo = [
        (0, {"words": [{"word": "Hello.", "start": 0.1, "end": 0.5}]}),
        (600, {"words": [{"word": "Again.", "start": 0.2, "end": 0.7}]}),
    ]
    json.dump(merge(demo), sys.stdout, indent=1)

A reconnect and chunk checklist

Whichever service you use, long audio turns into a bookkeeping problem. These habits keep it manageable.

  • Store the offset of every chunk next to its job id before you submit, not after the results come back.
  • Cut at the quietest point within a few seconds of the target length, so no word is split.
  • Submit chunks in parallel with a small cap, then sort by offset when merging.
  • Keep each chunk under 600 seconds and pass the real length as duration_seconds, which lets the reservation match the audio.
  • Re-run only the failed chunk rather than the whole file.

Limits and a caution

Each Sume result caps words[] at 20000 entries and reports words_truncated if it ever hits that, which a 600 second chunk stays well under. If you merge chunks yourself, keep the final merged list outside any single job's cap.

Not every case needs chunking. If you have a 12 minute file, two jobs of 6 minutes is simpler than a socket, and the per-job duration_seconds hint also makes the usage reservation match what you need. For live audio longer than an hour, plan the reconnect: the preview has no SLA, so the reconnect path is the part to test first.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume