Transcribe a 45-minute talk into a blog draft: five jobs, 45 cents

A 45-minute talk is five Sume STT jobs at 45 cents total. The slicing math, a stitching script, and what you still edit by hand before posting.

6 min readSume
All posts

A 45-minute talk costs 45 cents to transcribe on Sume: five STT 1.0 jobs, four covering 600 seconds each at 10 cents and one covering the last 300 seconds at 5 cents. Each job returns text and word timings, which you stitch into one transcript and edit into a post.

The job limit is the reason for the slicing. Sume STT accepts up to 600 seconds of audio per job, so 2,700 seconds needs five.

The slicing math

Sume bills STT at about one cent per audio minute, rounded up, with a one-cent minimum and a 10-cent maximum per job. Pass duration_seconds so the reservation matches the slice; leave it out and the reservation assumes one minute.

If the talk is a video on Sume's media host, cut the audio with audio detach: one POST /v1/audio-detach per slice with a range, $0.01 per job. Detach caps the source at 1,800 seconds and the output at 900 seconds, and it needs a media.sume.com video, so a 45-minute video does not fit as a single source. For a longer file, cut it with ffmpeg first and host each slice on a public HTTPS URL.

45-minute talk, Sume STT 1.0 cost by slice (read 2026-10-08)
SliceSecondsBilled
10 to 600$0.10
2600 to 1200$0.10
31200 to 1800$0.10
41800 to 2400$0.10
52400 to 2700$0.05
Total2700$0.45

Stitch the slices

Submit the five jobs, collect each result and offset the word times by the slice start so one timeline runs 0 to 2,700 seconds. The code below assumes you already have the five results as dicts with text and words.

Cut slices at silences if you can. A word split across a boundary will be transcribed as two fragments, and the join point is where you should read closely.

def stitch(results, slice_seconds=600):
    text, words = [], []
    for i, res in enumerate(results):
        shift = i * slice_seconds
        text.append(res["text"].strip())
        for w in res["words"]:
            words.append({**w,
                          "start": w["start"] + shift,
                          "end": w["end"] + shift})
    return " ".join(text), words

full_text, words = stitch([{"text": "a", "words": [{"word": "a", "start": 0.0, "end": 0.4}]},
                           {"text": "b", "words": [{"word": "b", "start": 0.0, "end": 0.3}]}])
print(full_text, words[-1]["start"])

What stays manual

The transcript is verbatim speech. A blog post needs headings, cut false starts and corrected names. STT 1.0 has no phrase list, so proper nouns may need fixing by hand; keep the speaker's slide deck open while you edit.

For comparison, Microsoft's launch posts list MAI-Transcribe-2 at $0.10 per hour as a limited-time offer to the end of the year, which would make this talk about 7.5 cents, and MAI-Transcribe-2-Streaming at $0.54 per hour, about 40.5 cents (read 2026-10-08). Sume's flat 10 cents for a full slice is higher per hour; its draw is one HTTP call per slice with no cloud account to set up.

From transcript to draft

Do not publish the raw text. Read it once for structure and mark where the speaker changed topic, which is usually where a heading goes. Use the word timings to look up the exact second of each topic change so you can link to the video at that point. Cut verbal tics and repeated starts, correct names against the slide deck and keep numbers exactly as spoken unless you can verify them.

A 45-minute talk at about 150 words a minute is roughly 6,750 words of transcript, which is a long post. Plan to publish a 1,200 to 1,800 word write-up and link the full transcript separately for readers who want it. The first pass takes minutes with the transcript; the editing is where the time goes, and the 45 cents is the least of it.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume