Transcribe a 2-hour recording: twelve 10-minute STT jobs, $1.24

A two-hour recording is twelve 10-minute Sume STT jobs at $0.01 a minute, cut with four Timeline audio split jobs: $1.24 at list rates, with the seam caveat.

4 min readSume
All posts

A two-hour recording takes twelve Sume speech-to-text jobs, because the pricing page lists a maximum of 10 minutes per transcription. At the listed $0.01 per audio minute that is $1.20 of transcription. Cutting the file into twelve windows with Timeline audio split adds four jobs at $0.01, so the list-rate total is $1.24. Both prices were read on 2026-10-03; confirm them live in GET /v1/catalog before budgeting.

How do I cut a long file into 10-minute windows?

Import the recording first: Sume's audio tools only take audio already on your workspace's media.sume.com host. If the recording is a video, run audio detach once and ask for 16000 Hz mono wav, which the docs call the speech-to-text shape.

Then call Timeline audio with operation: "split", the file url, and up to 20 ranges. The docs cap produced audio at 1800 seconds per job, so plan 30 minutes of windows per split job: three 600-second ranges each, four jobs for two hours. Every file returned has its own audio_url, ready to send to stt_create or POST /v1/stt-1.0/transcribe.

Window plan for a 2-hour file, rates read 2026-10-03 from Sume's pricing page
WindowStartsEndsSTT cost at $0.01 a minute
10:00:000:10:00$0.10
20:10:000:20:00$0.10
30:20:000:30:00$0.10
40:30:000:40:00$0.10
50:40:000:50:00$0.10
60:50:001:00:00$0.10
71:00:001:10:00$0.10
81:10:001:20:00$0.10
91:20:001:30:00$0.10
101:30:001:40:00$0.10
111:40:001:50:00$0.10
121:50:002:00:00$0.10
import math

TOTAL_SECONDS = 2 * 60 * 60
WINDOW = 600  # one STT request covers at most 10 minutes

windows = [
    (start, min(start + WINDOW, TOTAL_SECONDS))
    for start in range(0, TOTAL_SECONDS, WINDOW)
]
stt_cost = len(windows) * (WINDOW / 60) * 0.01
split_jobs = math.ceil(TOTAL_SECONDS / 1800)  # 1800 s of produced audio per split job
print(len(windows), "windows", round(stt_cost, 2), "USD STT")
print(split_jobs, "split jobs", round(split_jobs * 0.01, 2), "USD")
for start, end in windows[:3]:
    print({"start": start, "end": end})

What does the total include?

The $1.24 covers STT at list rate plus the four split jobs. It leaves out a detach job if your source is video (one more cent), any retries, and your own storage. Reservations are not the same as settled cost: the docs describe flat prices on the ffmpeg tools as list rates to confirm live, so treat the figure as a budget line, not an invoice.

What are the limits?

  • Cutting at fixed 10-minute marks can slice a word in half. Cut at silence if the transcript is for reading.
  • Each window is its own file, so each transcript's timings start at zero. Add the window start to every word time before you merge them.
  • I did not verify a file-size ceiling for STT input on a public page, so export compact audio and test one window first.
  • Speaker labels are not part of what the docs list for the transcript, which carries text, words[] and optional sentence segments[].

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume