Transcribe a 20-episode podcast back catalog: $5.60 on Sume

Twenty 25-minute video episodes cost $0.28 each to transcribe on Sume: three audio detach ranges at $0.01 and 25 STT minutes at $0.01, $5.60 in total.

4 min readSume
All posts

Transcribing 20 video podcast episodes of 25 minutes each costs $5.60 on Sume: $0.28 per episode, made of three audio detach jobs ($0.03) and 25 minutes of speech to text ($0.25). That makes a searchable archive for a show's back catalog a coffee-sized line item, and it is a sensible end-of-year task: roughly 90 days remain in 2026 for a best-of list and an updated website.

The constraint that shapes the plan is chunking. Speech to text accepts at most 600 seconds per job, and audio detach accepts a source of up to 1,800 seconds but returns at most 900 seconds of audio per job, so a 25-minute episode becomes three ranges.

Per-episode workflow

Import the episode video, then run audio detach three times with a range each (0 to 600, 600 to 1,200 and 1,200 to the end), mono at 16 kHz, which the docs name as the speech to text shape. Send each resulting audio URL to the transcription route with its duration and, if the show is in one language, a language code. Ask for sentence segmentation and you get timestamped sentences per chunk. Add each chunk's start offset (0, 600, 1,200) to its timestamps when you merge.

The transcript has no speaker labels, so host and guest attribution is a step you do yourself, usually from the show notes or the guest list. Each STT job is billed at $0.01 per audio minute, so the 10, 10 and 5 minute chunks come to $0.10, $0.10 and $0.05.

Back-catalog transcript cost by catalog size and episode length (read 2026-10-03)
CatalogPer-episode chunksPer episodeTotal
10 episodes of 25 min3 detach + 25 STT min$0.28$2.80
20 episodes of 25 min3 detach + 25 STT min$0.28$5.60
20 episodes of 30 min3 detach + 30 STT min$0.33$6.60
40 episodes of 25 min3 detach + 25 STT min$0.28$11.20

Limits that change the plan

Episodes longer than 30 minutes exceed the 1,800 second source cap on audio detach, so split those with your own editor before importing, or work from an audio file you host on a public HTTPS URL, which the transcription route accepts directly as long as each job stays under 600 seconds. Audio detach also needs a video source, so an audio-only feed has to be chunked outside Sume.

Do not treat the output as publishable copy. Names of guests, products and places are the usual errors in machine transcripts, and the request has no custom vocabulary field, so a find-and-replace pass over a glossary is the practical fix. Keep the raw transcript and the corrected one as separate files.

Once the transcripts exist, the same timestamps feed chapter lists and clip picks. A second pass over the best five moments per episode with video trim at $0.02 each is a separate budget, and it is cheaper to decide which episodes deserve it after reading the transcripts than before.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume