Transcribe a Zoom recording to text with an API, in 10-minute chunks
Turn a long meeting recording into text with Sume: detach the audio from the MP4, cut it into ranges under 10 minutes, and send each range to STT 1.0.

A Zoom recording is a video file, and Sume's STT 1.0 takes an audio URL of at most 10 minutes. So the route is three steps: pull the audio out of the MP4, split it into pieces under 600 seconds, and transcribe each piece. Each STT result returns text and words[] with start and end times in seconds from the start of that piece, so you add the piece's offset when you stitch the transcript together.
Step 1: detach the audio
Import the recording to your workspace first (POST /v1/media-imports), because audio detach only reads media.sume.com files. The source may be up to 1800 seconds and one output up to 900 seconds, so a longer recording needs a range. The docs name 16000 Hz mono as the STT shape.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: zoom-detach-001" \
-d '{"video_url": "'"$MEETING_URL"'", "format": "wav", "channels": "mono", "sample_rate": 16000, "range": {"start": 0, "end": 600}}'Step 2: one STT job per piece
Repeat with range windows of 600 seconds, or detach once and split with timeline audio operation: "split", which takes 1 to 20 ranges. A recording with no audio track fails with detach_source_has_no_audio.
Send each audio_url to POST /v1/stt-1.0/transcribe. Pass duration_seconds (1 to 600). If you omit it, Sume reserves one minute, so a ten-minute piece reserves too little. Omit language_code to let the model detect the language, or pass a hint such as en. Add segmentation: {"mode": "sentence"} to also get sentence segments. metadata is an optional free-form object.
The limits and rates that decide the plan, read 2026-10-06:
| Stage | Limit | Public rate |
|---|---|---|
| Audio detach | Source 1800 s, output 900 s | $0.01 per job |
| Timeline audio split | 1 to 20 ranges per job | $0.01 per job |
| STT 1.0 | 600 s per audio file | $0.01 per audio minute |
Stitching the chunks back together
Each STT result times its words from the start of its own piece. Add the piece's start offset, for example 600 seconds for the second ten-minute range, to every start and end before you merge. Without that, the second half of the transcript points at the first ten minutes.
Cut on silence where you can. A range boundary in the middle of a word loses it in both pieces, so start each range a second early and drop the overlap when you merge. Keep the job ids in a list, since the result is read by job id and not by file name, and re-read any piece that fails without resubmitting the others.
For a 30-minute meeting that is one detach job per 10-minute range or one detach plus one split, then three STT jobs. At the listed rates that is a few cents in total. Confirm the live prices in GET /v1/catalog before you budget. If you only need words as they are spoken in a live call, a batch job is the wrong tool, but for a recording it gives you timestamps you can reuse for chapters or captions.
Sources
Related posts
More in Developers
- Trim silence from a voiceover using STT word timings (Python)
Find the dead air in a voiceover from Sume STT words[] timings: list gaps over a threshold in Python, then cut with timeline audio ranges.
- Text-to-speech API in Node: submit, poll and save the MP3
A Node 18+ fetch example for the Sume TTS Router: submit with an Idempotency-Key, poll status_url, read the audio artifact and save an MP3 to disk.
- TTS model list API: read the catalog before you hardcode a Sonic id
Sume's TTS Router lists its models at GET /v1/tts-router/models. Read it, pin sonic-3.6, and treat sonic-preview as a beta channel that can change.
- tts_sentence_selection_invalid 422 on Sume TTS: what triggers it
Sume TTS returns 422 tts_sentence_selection_invalid for gaps, repeated jobs, unfinished jobs and partial coverage. Each cause and its fix.
Written by Sume