Transcribe a Zoom recording to text with an API, in 10-minute chunks

Turn a long meeting recording into text with Sume: detach the audio from the MP4, cut it into ranges under 10 minutes, and send each range to STT 1.0.

5 min readSume
All posts

A Zoom recording is a video file, and Sume's STT 1.0 takes an audio URL of at most 10 minutes. So the route is three steps: pull the audio out of the MP4, split it into pieces under 600 seconds, and transcribe each piece. Each STT result returns text and words[] with start and end times in seconds from the start of that piece, so you add the piece's offset when you stitch the transcript together.

Step 1: detach the audio

Import the recording to your workspace first (POST /v1/media-imports), because audio detach only reads media.sume.com files. The source may be up to 1800 seconds and one output up to 900 seconds, so a longer recording needs a range. The docs name 16000 Hz mono as the STT shape.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: zoom-detach-001" \
  -d '{"video_url": "'"$MEETING_URL"'", "format": "wav", "channels": "mono", "sample_rate": 16000, "range": {"start": 0, "end": 600}}'

Step 2: one STT job per piece

Repeat with range windows of 600 seconds, or detach once and split with timeline audio operation: "split", which takes 1 to 20 ranges. A recording with no audio track fails with detach_source_has_no_audio.

Send each audio_url to POST /v1/stt-1.0/transcribe. Pass duration_seconds (1 to 600). If you omit it, Sume reserves one minute, so a ten-minute piece reserves too little. Omit language_code to let the model detect the language, or pass a hint such as en. Add segmentation: {"mode": "sentence"} to also get sentence segments. metadata is an optional free-form object.

The limits and rates that decide the plan, read 2026-10-06:

Zoom recording pipeline, read 2026-10-06
StageLimitPublic rate
Audio detachSource 1800 s, output 900 s$0.01 per job
Timeline audio split1 to 20 ranges per job$0.01 per job
STT 1.0600 s per audio file$0.01 per audio minute

Stitching the chunks back together

Each STT result times its words from the start of its own piece. Add the piece's start offset, for example 600 seconds for the second ten-minute range, to every start and end before you merge. Without that, the second half of the transcript points at the first ten minutes.

Cut on silence where you can. A range boundary in the middle of a word loses it in both pieces, so start each range a second early and drop the overlap when you merge. Keep the job ids in a list, since the result is read by job id and not by file name, and re-read any piece that fails without resubmitting the others.

For a 30-minute meeting that is one detach job per 10-minute range or one detach plus one split, then three STT jobs. At the listed rates that is a few cents in total. Confirm the live prices in GET /v1/catalog before you budget. If you only need words as they are spoken in a live call, a batch job is the wrong tool, but for a recording it gives you timestamps you can reuse for chapters or captions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume