Show notes from a video: timestamped sentences via Sume STT

Detach a 600-second mono wav, run STT with sentence segmentation, and get timestamped sentences for show notes. Each 10 minutes costs $0.11.

3 min readSume
All posts

Detach a mono 16 kHz wav of up to 600 seconds, send it to POST /v1/stt-1.0/transcribe with segmentation.mode: "sentence", and read segments[] for timestamped sentences. Each 10-minute chunk costs $0.01 for the detach plus $0.10 for the transcript.

The two calls

Audio detach is $0.01 per job and accepts a range; channels: "mono" with sample_rate: 16000 is the shape the docs name as the STT input. The STT job takes the resulting audio_url, a duration_seconds between 1 and 600, and an optional language_code.

Cost of show notes by episode length (read 2026-10-03)
Video lengthChunksDetachSTTTotal
10 minutes1$0.01$0.10$0.11
30 minutes3$0.03$0.30$0.33
60 minutes6$0.06$0.60$0.66

Reading the result

Sentence segments are gapless and ordered, so segments[i].end equals segments[i+1].start. They are time ranges over the submitted audio, and no sliced files are produced. For chunks after the first, add the chunk start to each time before you merge. The result has no speaker labels, so attribute quotes yourself.

From segments to notes

Sume gives the timestamped text; the summarising is yours, whether by hand or with a language model. Pick the sentences that open a topic and link them to the time. The chapters post covers turning timestamps into a chapter list.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume