Stitch three STT chunks: add each range start to word times
Sume STT times count from each chunk's own start, so add the chunk's range start to every word and sentence. A short Python function that does it.

Add the start of each chunk's range, in seconds, to every start and end in that chunk's result, then concatenate the texts and word lists in order. Sume STT reports each time as an offset in seconds from the start of the audio file it read, so chunk two's first word reads 0.2 even if it is spoken at 600.2 seconds of the lecture.
One STT job reads at most 600 seconds, so a longer recording must be cut and the pieces transcribed separately. This post shows the joining step. Microsoft's new streaming model, launched 2026-10-01 (Microsoft AI, read 2026-10-05), is aimed at live audio, and the chunk join is the price of the file route on Sume.
What each result gives you
The result fields that matter are text, words[] and, when you submitted segmentation: {"mode": "sentence"}, segments[]. Each word has word, start and end, and a type when the provider supplies one. Each segment has index, text, start, end and duration_seconds. Segments are gapless within one chunk: each ends exactly where the next begins.
Segment index restarts at zero in every chunk, so after the join you must renumber it. The function below does that as its last step. Duration needs no change, because it is a difference of two times.
The function
The function takes a list of pairs, the range start you cut at and the STT result for that chunk. It sorts by the range start, so the order you finished the jobs in does not matter. It also copies each entry instead of editing your stored results. The demo at the bottom runs with no network and no key, so you can test it on its own.
The offset you add must be the real start of the cut file within the original recording. If you cut with Audio detach, that is the range.start you sent. Take it from your own submit record, not from the result.
def stitch(chunks):
"""chunks: list of (range_start_seconds, stt_result_dict)."""
words, sentences, texts = [], [], []
for start, res in sorted(chunks, key=lambda c: c[0]):
texts.append(res.get("text", "").strip())
for w in res.get("words", []):
w = dict(w)
for k in ("start", "end"):
if k in w:
w[k] = round(w[k] + start, 3)
words.append(w)
for seg in res.get("segments", []):
seg = dict(seg)
seg["start"] = round(seg["start"] + start, 3)
seg["end"] = round(seg["end"] + start, 3)
sentences.append(seg)
for i, seg in enumerate(sentences):
seg["index"] = i
return " ".join(texts), words, sentences
a = {"text": "Hello there.", "words": [{"word": "Hello", "start": 0.1, "end": 0.5}],
"segments": [{"index": 0, "text": "Hello there.", "start": 0.0, "end": 1.2}]}
b = {"text": "Next part.", "words": [{"word": "Next", "start": 0.2, "end": 0.6}],
"segments": [{"index": 0, "text": "Next part.", "start": 0.0, "end": 1.0}]}
text, words, sents = stitch([(600, b), (0, a)])
print(text, words[1]["start"], sents[1]["index"], sents[1]["start"])Overlaps and boundaries
If you cut chunks with a few seconds of overlap so no word is split, you will see words twice around each boundary. After adding offsets, drop a word from the later chunk when its start is earlier than the end of the last word kept from the earlier chunk. Do the same for sentences. The offset step stays the same.
A cut in mid-word gives two half words. An overlap of two or three seconds, plus the rule above, is a good fix, and a pause-aware cut is better still if your audio has clear gaps.
Edge cases the result can hold
A result is capped at 20,000 words. A 600-second transcript stays well under that, and if the cap is ever reached the result also carries words_truncated and words_total, so check for them before you trust a joined transcript. The cap applies per chunk, which is another reason to keep chunks short.
A word can come back without start or end. The function leaves those untouched instead of adding an offset to a missing value, which is why it tests for each key.
A quick check
Check the join with two numbers. The last word's end after the join should be close to the length of the recording, and the sum of the chunk durations should equal it. If the last word ends far earlier, a chunk's job failed and you joined fewer pieces than you thought.
Keep the joined result next to the chunk job ids. A transcript that can be traced to three job results can be rebuilt in a second, and one that cannot is a file nobody dares to touch.
Finally, round only at the end. The function rounds to three decimals, which is finer than a caption frame, so it never moves a word by a visible amount. Feed the joined words into your caption or search code in place of any single chunk's list, and everything downstream keeps working on lecture time.
Sources
Related posts
More in Developers
- Stop an Omni draft batch at $5: sum usage.cost from each poll
A Python loop that submits 360p Omni drafts one at a time, adds the Sume usage.cost of each finished job, and stops before the next one would cross a $5 cap.
- Can Strands Decider 2B pick which Sume MCP tool to call?
Strands Decider 2B scores options you give it. Build them from Sume's tools_list, then check the pick with tools_schema and dry_run before any paid call.
- Stripe allows 16 webhook endpoints; Sume takes a webhook URL per job
Stripe registers up to 16 endpoints. Sume has no registry: each job or Format run item carries its own public HTTPS webhook_url, signed with one secret.
- Stripe events arrive out of order; Sume sends only terminal job events
Stripe does not order events and says to dedupe on event ID. Sume sends only terminal job events keyed by job_id, so order rarely matters. Fetch state.
Written by Sume