Transcribe a 2-hour recording: twelve 10-minute STT jobs, $1.24
A two-hour recording is twelve 10-minute Sume STT jobs at $0.01 a minute, cut with four Timeline audio split jobs: $1.24 at list rates, with the seam caveat.

A two-hour recording takes twelve Sume speech-to-text jobs, because the pricing page lists a maximum of 10 minutes per transcription. At the listed $0.01 per audio minute that is $1.20 of transcription. Cutting the file into twelve windows with Timeline audio split adds four jobs at $0.01, so the list-rate total is $1.24. Both prices were read on 2026-10-03; confirm them live in GET /v1/catalog before budgeting.
How do I cut a long file into 10-minute windows?
Import the recording first: Sume's audio tools only take audio already on your workspace's media.sume.com host. If the recording is a video, run audio detach once and ask for 16000 Hz mono wav, which the docs call the speech-to-text shape.
Then call Timeline audio with operation: "split", the file url, and up to 20 ranges. The docs cap produced audio at 1800 seconds per job, so plan 30 minutes of windows per split job: three 600-second ranges each, four jobs for two hours. Every file returned has its own audio_url, ready to send to stt_create or POST /v1/stt-1.0/transcribe.
| Window | Starts | Ends | STT cost at $0.01 a minute |
|---|---|---|---|
| 1 | 0:00:00 | 0:10:00 | $0.10 |
| 2 | 0:10:00 | 0:20:00 | $0.10 |
| 3 | 0:20:00 | 0:30:00 | $0.10 |
| 4 | 0:30:00 | 0:40:00 | $0.10 |
| 5 | 0:40:00 | 0:50:00 | $0.10 |
| 6 | 0:50:00 | 1:00:00 | $0.10 |
| 7 | 1:00:00 | 1:10:00 | $0.10 |
| 8 | 1:10:00 | 1:20:00 | $0.10 |
| 9 | 1:20:00 | 1:30:00 | $0.10 |
| 10 | 1:30:00 | 1:40:00 | $0.10 |
| 11 | 1:40:00 | 1:50:00 | $0.10 |
| 12 | 1:50:00 | 2:00:00 | $0.10 |
import math
TOTAL_SECONDS = 2 * 60 * 60
WINDOW = 600 # one STT request covers at most 10 minutes
windows = [
(start, min(start + WINDOW, TOTAL_SECONDS))
for start in range(0, TOTAL_SECONDS, WINDOW)
]
stt_cost = len(windows) * (WINDOW / 60) * 0.01
split_jobs = math.ceil(TOTAL_SECONDS / 1800) # 1800 s of produced audio per split job
print(len(windows), "windows", round(stt_cost, 2), "USD STT")
print(split_jobs, "split jobs", round(split_jobs * 0.01, 2), "USD")
for start, end in windows[:3]:
print({"start": start, "end": end})
What does the total include?
The $1.24 covers STT at list rate plus the four split jobs. It leaves out a detach job if your source is video (one more cent), any retries, and your own storage. Reservations are not the same as settled cost: the docs describe flat prices on the ffmpeg tools as list rates to confirm live, so treat the figure as a budget line, not an invoice.
What are the limits?
- Cutting at fixed 10-minute marks can slice a word in half. Cut at silence if the transcript is for reading.
- Each window is its own file, so each transcript's timings start at zero. Add the window start to every word time before you merge them.
- I did not verify a file-size ceiling for STT input on a public page, so export compact audio and test one window first.
- Speaker labels are not part of what the docs list for the transcript, which carries text,
words[]and optional sentencesegments[].
Sources
Related posts
More in Use cases
- Trendyol image size 1200x1800 and what to ask Sume for
Trendyol's Product Create v2 API doc: 1200x1800 px at 96 dpi, up to 8 images per barcode, https URLs. The 2:3 Sume request that fits.
- Trim before you recast: pay for the seconds the swap needs
Recast bills the whole source length. Cut a 30-second clip to the 12 seconds that matter with POST /v1/video-trim and the 768p price drops from $11.25 to $4.50.
- Trim a UGC ad down to its hook: frame-accurate cut for $0.02
Sume video trim cuts a start/end range from a hosted clip into a new MP4 for $0.02 a job. Exact vs keyframe precision, the 0.2-900 s duration, and import first.
- Turkey trot registration video: 3 versions of one 5K clip for $1.23
Thanksgiving is Thursday 26 November 2026. Generate one 5-second race clip and burn early, last-call and charity lines on it for about $1.23 on Sume.
Written by Sume